SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Jiayuan Gu, Zhigang Wang, Yan Fang Ding 等 11 位
Fudan University Shanghai Artificial Intelligence Laboratory Shanghai Jiao Tong University Zhejiang University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
and real-world setups, where the pre-learned action grids are re-discretized to capture robot-specific spatial action movements of new setups.The superior results from extensive evaluations demonstrate the exceptional in-distribution generalization and out-of-distribution adaptation capability, highlighting the crucial benefit of the proposed spatial-aware representations for generalist robot policy learning.All the details and codes are opensourced.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Geographic Information Systems Studies · Human Pose and Action Recognition
参考文献 0
引用本文 41
按被引量排序,此处列出前 3 条