ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
Liang Heng, Haoran Geng, Kaifeng Zhang, Pieter Abbeel, Jitendra Malik
Peking University University of California, Berkeley
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Dexterous manipulation is a cornerstone capability for robotic systems aiming to interact with the physical world in a human-like manner.Although vision-based methods have advanced rapidly, tactile sensing remains crucial for fine-grained control-particularly in unstructured or visually occluded settings.We present ViTacFormer, a representationlearning approach that couples a cross-attention encoder to fuse high-resolution vision and touch with an autoregressive tactileprediction head that anticipates future contact signals.Building on this architecture, we devise an easy-to-challenging curriculum that steadily refines the visual-tactile latent space, boosting both accuracy and robustness.The learned cross-modal representation drives imitation learning for multi-fingered hands, enabling precise and adaptive manipulation.Across a suite of challenging real-world benchmarks, our method achieves approximately 50% higher success rates than prior state-of-the-art systems.To our knowledge, it is also the first to autonomously complete longhorizon dexterous manipulation tasks that demand highly precise control with an anthropomorphic hand-successfully executing up to 11 sequential stages and sustaining continuous operation for 2.5 minutes.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Advanced Sensor and Energy Harvesting Materials
Robot Manipulation and Learning · Tactile and Sensory Interactions
参考文献 0
引用本文 1
按被引量排序,此处列出前 3 条