ModDrop: Adaptive Multi-Modal Gesture Recognition
Natalia V. Neverova, Christian Wolf, Graham Taylor, Florian Nebout
Centre National de la Recherche Scientifique Université de Lyon Laboratoire d'Informatique en Images et Systèmes d'Information Institut National des Sciences Appliquées de Lyon
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
We present a method for gesture detection and localisation based on multi-scale and multi-modal deep learning. Each visual modality captures spatial information at a particular spatial scale (such as motion of the upper body or a hand), and the whole system operates at three temporal scales. Key to our technique is a training strategy which exploits: i) careful initialization of individual modalities; and ii) gradual fusion involving random dropping of separate channels (dubbed ModDrop) for learning cross-modality correlations while preserving uniqueness of each modality-specific representation. We present experiments on the ChaLearn 2014 Looking at People Challenge gesture recognition track, in which we placed first out of 17 teams. Fusing multiple modalities at several spatial and temporal scales leads to a significant increase in recognition rates, allowing the model to compensate for errors of the individual classifiers as well as noise in the separate channels. Furthermore, the proposed ModDrop training technique ensures robustness of the classifier to missing signals in one or several channels to produce meaningful predictions from any number of available modalities. In addition, we demonstrate the applicability of the proposed fusion scheme to modalities of arbitrary nature by experiments on the same dataset augmented with audio.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIHand Gesture Recognition Systems
Human Pose and Action Recognition · Interactive and Immersive Displays
参考文献 77
此处列出前 3 条
引用本文 412
按被引量排序,此处列出前 3 条