PointVLA: Injecting the 3D World Into Vision-Language-Action Models
Chengmeng Li, Junjie Wen, Yaxin Peng, Yan Peng, Yichen Zhu
Shanghai University Midea Group (China)
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning critical for real-world interaction. Retraining these models with 3D data is computationally prohibitive, while discarding existing 2D datasets wastes valuable resources. To bridge this gap, we propose PointVLA, a framework that enhances pre-trained VLAs with point cloud inputs without requiring retraining. Our method freezes the vanilla action expert and injects 3D features via alightweight modular block. To identify the most effective way of integrating point cloud representations, we conduct a skip-block analysis to pinpoint less useful blocks in the vanilla action expert, ensuring that 3D features are injected only into these blocks—minimizing disruption to pre-trained representations. Extensive experiments demonstrate that PointVLA outperforms state-of-the-art 2D imitation learning methods, such as OpenVLA, Diffusion Policy and DexVLA, across both simulated and real-world robotic tasks. Specifically, we highlight several key advantages of PointVLA enabled by point cloud integration: (1)Few-shotmulti-tasking, where PointVLA successfully performs four different tasks using only 20 demonstrations each; (2)Real-vs-photo discrimination, where PointVLA distinguishes real objects from their images, leveraging 3D world knowledge to improve safety and reliability; (3)Height adaptability, where unlike conventional 2D imitation learning methods, PointVLA enables robots to adapt to objects at varying table heights that were unseen in training data. Furthermore, PointVLA achieves strong performance in long-horizon tasks, such as picking and packing objects from a moving conveyor belt, showcasing its ability to generalize across complex, dynamic environments.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Robot Manipulation and Learning · Advanced Neural Network Applications
参考文献 21
此处列出前 3 条
引用本文 15
按被引量排序,此处列出前 3 条