MDFDET: A Multi-Modal Dynamic Fusion Algorithm for RGB-Infrared Object Detection
Ji Feng, Bingtao Liu, Cuijin Li, Chuan Huang, J Liu
Chongqing Normal University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
In low-illumination traffic scenarios, multimodal object detection encounters challenges such as weak cross-modal interaction, insufficient detail and boundary fidelity in upsampling, and limited global selection with noise suppression, leading to suboptimal multisource complementarity. This paper proposes a multi-modal dynamic perception fusion detection algorithm based on visible and infrared modalities. A cross-layer RGB-Infrared mid-level fusion backbone is designed to mitigate single-modal degradation and domain shift, enabling deeper interaction. To enhance representation, a C3M (Cross-Stage Convolutional-Mixed Multi-scale Attention Module) addresses scale and semantic deficiencies, while DySample is embedded in the neck for adaptive upsampling, preserving detail fidelity. Furthermore, a GSCSA (Grouped Spatial-Channel Synergistic Attention) module is added before the detection head to suppress semantic mismatch and environmental noise. Extensive experiments on public traffic datasets validate the framework: it improves mAP50:95 by 8.3% on M3FD dataset and by 5.6% on DroneVehicle dataset compared with state-of-the-art methods, achieving superior accuracy without compromising efficiency.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Infrared Target Detection Methodologies
Advanced Neural Network Applications · Medical Image Segmentation Techniques
参考文献 19
此处列出前 3 条