Cross-Modal Alignment Enhancement Network for Text-to-Image Person Re-Identification
Di He, Xinshan Zhu, Bin Li, Shenglu Yue, Zhong Zhang
Tianjin University Tianjin Normal University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
With the proliferation of visual Internet of Things (VIoT), the intelligent processing of imagery and textual data to accomplish text-to-image person re-identification (Re-ID) constitutes a formidable challenge within VIoT, where text-to-image person Re-ID endeavors to execute image retrieval across disparate visual sensors through designated text query. The existing works mainly focus on learning feature alignment between texts and images from global and local perspectives. However, these methods exhibit two fundamental limitations: 1) neglecting the relationship between different kinds of positive pairs (i.e., original positive pairs and weak positive pairs) during global feature alignment; 2) employing horizontally partitioned stripe parts of pedestrian images for local feature alignment constrains adaptability. To address these issues, we propose a novel cross-modal Alignment Enhancement Network (AENet), which effectively align cross-modal features by modeling the relationships among different positive pairs and enabling flexible local correspondences. Specifically, we design a Hybrid Positive Pair Learning (HPPL) strategy to fully exploit the complementary semantic information in original and weak positive pairs through adaptive fusion, thereby enhancing global feature alignment. Regarding local feature alignment, we develop a Local Aggregation and Dynamic Matching (LADM) module to learn fine-grained correspondences through generating representative cross-modal local aggregation features and implementing dynamic matching mechanism. Correspondingly, we propose a Token Reconstruction Loss (TRL) to regulate the local aggregation process, thereby ensuring the quality of local aggregation features. Extensive experiments on three publicly available benchmarks verify the effectiveness and superiority of the proposed method, exemplified by achieving 76.48%, 68.34%, and 66.70% Rank-1 accuracy on CUHK-PEDES, ICFG-PEDES, and RSTPReid, respectively.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIVideo Surveillance and Tracking Methods
Advanced Neural Network Applications · Multimodal Machine Learning Applications
参考文献 66
此处列出前 3 条
引用本文 2
按被引量排序,此处列出前 3 条