emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, Xie Chen
Shanghai Jiao Tong University Fudan University Chinese University of Hong Kong
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
We propose emotion2vec, a universal speech emotion representation model.emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and framelevel loss during pre-training.emotion2vec outperforms state-of-the-art pre-trained universal models and emotion specialist models by only training linear layers for the speech emotion recognition task on the mainstream IEMOCAP dataset.In addition, emotion2vec shows consistent improvements among 10 different languages of speech emotion recognition datasets.emotion2vec also shows excellent results on other emotion tasks, such as song emotion recognition, emotion prediction in conversation, and sentiment analysis.Comparison experiments, ablation experiments, and visualization comprehensively demonstrate the universal capability of the proposed emotion2vec.To the best of our knowledge, emotion2vec is the first universal representation model in various emotion-related tasks, filling a gap in the field.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AISpeech Recognition and Synthesis
Emotion and Mood Recognition · Speech and dialogue systems
参考文献 0
引用本文 130
按被引量排序,此处列出前 3 条