Transformer-based Feature Reconstruction Network for Robust Multimodal Sentiment Analysis
Ziqi Yuan, Wei Li, Hua Xu, Wenmeng Yu
Beijing National Research Center for Information Science and Technology Tsinghua University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Improving robustness against data missing has become one of the core challenges in Multimodal Sentiment Analysis (MSA), which aims to judge speaker sentiments from the language, visual, and acoustic signals. In the current research, translation-based methods and tensor regularization methods are proposed for MSA with incomplete modality features. However, both of them fail to cope with random modality feature missing in non-aligned sequences. In this paper, a transformer-based feature reconstruction network (TFR-Net) is proposed to improve the robustness of models for the random missing in non-aligned modality sequences. First, intra-modal and inter-modal attention-based extractors are adopted to learn robust representations for each element in modality sequences. Then, a reconstruction module is proposed to generate the missing modality features. With the supervision of SmoothL1Loss between generated and complete sequences, TFR-Net is expected to learn semantic-level features corresponding to missing features. Extensive experiments on two public benchmark datasets show that our model achieves good results against data missing across various missing modality combinations and various missing degrees.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AISpeech and Audio Processing
Music and Audio Processing · Multimodal Machine Learning Applications
参考文献 36
此处列出前 3 条
引用本文 187
按被引量排序,此处列出前 3 条