开发者可以使用平台所提供的脚本及各数据标准剪切集来开发及评估语音情绪模型;得到准确率后,可以提出要求将所得到的准确率纳入排行榜中;穿透此平台,开发者可以与平台提供的结果做比较,也可以利用平台做出有趣的分析。

Conference: 2024 IEEE Spoken Language Technology (SLT) Workshop

Paper:https://reurl.cc/KdrG2y

Code: https://github.com/EMOsuperb/EMO-SUPERB-submission

Website: https://emosuperb.github.io/


语音情绪识别是人机交互中至关重要的任务之一,最近取得了很大进展。然而,在语音情绪辨识任务中,仍然存在两个未解决的问题:
  1. 重现性的困难:知名情绪数据库IEMOCAP的作者发现有80.77%的论文无法重现[1]。
  2. 官方未提供数据集的官方切分方式导致了训练、开发和测试集的标准化缺失,各研究团队采取自己的数据切割方式,导致了比较的不公平性,甚至可能引发数据泄露问题。在情绪识别的训练数据中,通常是对话形式,例如一个对话中有说话人A和说话人B。在切分这段对话并分出说话人A和说话人B的语音片段时,很容易出现说话人A的语音片段中夹杂着说话人B的语音。许多数据集仅仅是根据说话人进行简单的数据切分,这样很可能会导致将说话人A切分到训练数据中,说话人B切分到测试数据中。这就意味着模型在训练阶段已经接触到了说话人B的数据。

针对上述的挑战,我们提出方法,我们基于SUPERB[3]准备了EMO-SUPERB的平台,此平台包含6个公开资料库,含有英文与中文情绪语料库,并针对一些没有提供标准切割集的情绪资料库(例如,IEMOCAP),提供我们的切割方法,并提供所有资料库的标记以及训练脚本,让大家可以较无痛实作和重现语音情绪辨识的任务,而我们的网站,也提供了排行榜,欢迎大家参与。


此平台目前纳入六个不同的公开语料库,以及16个SSLMs,欢迎开发者增加资料库及模型。


图三、排行榜

此表格收录目前16个SSLMs以及FBANK对于六个情绪资料库中的九个不同情况的结果,评价指标用的是macro-F1 scores。


图四、雷达图

我们加入了一种State-of-the-art (SOTA) SER model[2],与EMO-SUPERB提供的模型做比较。


目前有包含的情绪资料库

  • MSP-IMPROV (IMPROV)[4]

  • CREMA-D[5]

  • MSP-PODCAST (POD) v1.11[6]

  • BIIC-PODCAST (B-POD) v1.01[7]

  • IEMOCAP[8]

  • NNIME[9]


相关研究

我们定义了语音情绪识别任务为多标籤的任务,一句音档可能包含一个或多个的情绪,其中的比较,我们在这一篇研究做深度的比较[9]。


参考资料:

[1]Antoniou, N., Katsamanis, A., Giannakopoulos, T., & Narayanan, S. (2023, June). Designing and Evaluating Speech Emotion Recognition Systems: A reality check case study with IEMOCAP. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 1-5). IEEE.

[2] Wagner, J., Triantafyllopoulos, A., Wierstorf, H., Schmitt, M., Burkhardt, F., Eyben, F., & Schuller, B. W. (2023). Dawn of the transformer era in speech emotion recognition: closing the valence gap. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), 10745-10759.

[3]Yang, S.-w., Chi, P.-H., Chuang, Y.-S., Lai, C.-I.J., Lakhotia, K., Lin, Y.Y., Liu, A.T., Shi, J., Chang, X., Lin, G.-T., Huang, T.-H., Tseng, W.-C., Lee, K.-t., Liu, D.-R., Huang, Z., Dong, S., Li, S.-W., Watanabe, S., Mohamed, A., Lee, H.-y. (2021) SUPERB: Speech Processing Universal PERformance Benchmark. Proc. Interspeech 2021, 1194-1198, doi: 10.21437/Interspeech.2021-1775

[3] Busso, C., Parthasarathy, S., Burmania, A., AbdelWahab, M., Sadoughi, N., & Provost, E. M. (2016). MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception. IEEE Transactions on Affective Computing, 8(1), 67-80.

[4] Cao, H., Cooper, D. G., Keutmann, M. K., Gur, R. C., Nenkova, A., & Verma, R. (2014). Crema-d: Crowd-sourced emotional multimodal actors dataset. IEEE transactions on affective computing, 5(4), 377-390.

[5] Lotfian, R., & Busso, C. (2017). Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings. IEEE Transactions on Affective Computing, 10(4), 471-483.

[6] Upadhyay, S. G., Chien, W. S., Su, B. H., Goncalves, L., Wu, Y. T., Salman, A. N., ... & Lee, C. C. (2023, September). An intelligent infrastructure toward large scale naturalistic affective speech corpora collection. In 2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII) (pp. 1-8). IEEE.

[7] Busso, C., Bulut, M., Lee, C. C., Kazemzadeh, A., Mower, E., Kim, S., ... & Narayanan, S. S. (2008). IEMOCAP: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42, 335-359.

[8] Chou, H. C., Lin, W. C., Chang, L. C., Li, C. C., Ma, H. P., & Lee, C. C. (2017, October). NNIME: The NTHU-NTUA Chinese interactive multimodal emotion corpus. In 2017 Seventh international conference on affective computing and intelligent interaction (ACII) (pp. 292-298). IEEE.

[9]Chou, H. C., Wu, H., Goncalves, L., Leem, S. G., Salman, A., Busso, C., ... & Lee, C. C.Embracing Ambiguity and Subjectivity Using the All-inclusive Aggregation Rule for Evaluating Multi-Label Speech Emotion Recognition Systems. In 2024 IEEE SLT.