Multi-Modal Mutual Attention and Iterative Interaction for Referring Image Segmentation
Chang Liu, Henghui Ding, Yulun Zhang, Xudong Jiang
Nanyang Technological University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
We address the problem of referring image segmentation that aims to generate a mask for the object specified by a natural language expression. Many recent works utilize Transformer to extract features for the target object by aggregating the attended visual regions. However, the generic attention mechanism in Transformer only uses the language input for attention weight calculation, which does not explicitly fuse language features in its output. Thus, its output feature is dominated by vision information, which limits the model to comprehensively understand the multi-modal information, and brings uncertainty for the subsequent mask decoder to extract the output mask. To address this issue, we propose Multi-Modal Mutual Attention (M3Att) and Multi-Modal Mutual Decoder (M3Dec) that better fuse information from the two input modalities. Based on M3Dec, we further propose Iterative Multi-modal Interaction (IMI) to allow continuous and in-depth interactions between language and vision features. Furthermore, we introduce Language Feature Reconstruction (LFR) to prevent the language information from being lost or distorted in the extracted feature. Extensive experiments show that our proposed approach significantly improves the baseline and outperforms state-of-the-art referring image segmentation methods on RefCOCO series datasets consistently.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Domain Adaptation and Few-Shot Learning · Topic Modeling
参考文献 84
此处列出前 3 条
引用本文 75
按被引量排序,此处列出前 3 条