Joint visual template and natural language for robust visual tracking
Jingchao Wang, Huanlong Zhang, Jianwei Zhang
Zhengzhou University of Light Industry
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
At present, the target of interest in visual tracking is given in the form of a bounding box. Due to the randomness of the target shape, the bounding box may contain a lot of non‐target information. When encountering complex tracking scenarios, the performance of the tracker reduces severely. To address this problem, in this letter, the authors propose a novel tracking framework based on the joint of the visual template and natural language (VNTrack) to alleviate the impact of bounding box ambiguity. Specifically, the authors first use a pre‐trained language model to extract the features of the language description of the target. Then, a feature alignment module is designed to align and enhance the visual template feature and natural language feature. In addition, the authors design a multimodal query module to fuse the visual template, natural language, and search region information. Experimental results over tracking benchmarks with language annotations show that the proposed VNTrack is competitive among the state‐of‐the‐art trackers.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIVideo Surveillance and Tracking Methods
Advanced Image and Video Retrieval Techniques · Human Pose and Action Recognition
参考文献 22
此处列出前 3 条
引用本文 4
按被引量排序,此处列出前 3 条