Using YOLOv7 Combined with Visual Attention Mechanism to Achieve Real-time Tracking of Passenger Gaze Targets in Urban Rail Interactive Guidance Interface
Lu Zhai, X. H. Wan, Y. M. Zhao, L. F. Li, L. X. Wu
Beijing Jiaotong University Beijing Urban Construction Design & Development Group (China) City University of Macau
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Real-time gaze target tracking in urban rail interactive guidance interfaces faces significant challenges due to multitarget occlusion, limited computational resources, and the need for low-latency visual perception. This study proposes a lightweight tracking framework that integrates an improved YOLOv7 detector with the Convolutional Block Attention Module (CBAM) and an enhanced ByteTrack algorithm to achieve efficient and accurate passenger gaze target localization. MobileNetV3 is adopted as the backbone network to reduce computational complexity, while BiFPN and Gaussian Soft-NMS improve multi-scale feature fusion and detection robustness. MediaPipe and PnP-based gaze estimation are combined with an NSA Kalman filter and GIoU matching strategy to enhance temporal consistency and tracking stability under complex scenarios. Experimental results demonstrate that the proposed framework achieves a mAP@0.5 of 96.3%, a MOTP of 85.7%, and 32.2 FPS on embedded platforms, maintaining reliable performance under severe occlusion conditions. Beyond intelligent transportation applications, the proposed lightweight visual perception and attention modeling framework provides useful methodological references for adaptive sensing, edge intelligence, and real-time signal interpretation in electromagnetic wave propagation environments and antenna-assisted human– machine interaction systems.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIGaze Tracking and Assistive Technology
Visual Attention and Saliency Detection · Hand Gesture Recognition Systems