Any-point Trajectory Modeling for Policy Learning
Chuan Wen, Xingyu Lin, John So, Kai Chen, Qi Dou, Yang Gao, Pieter Abbeel
内容与影响
Learning from demonstration is a powerful method for teaching robots new skills, and having more demonstration data often improves policy learning.However, the high cost of collecting demonstration data is a significant bottleneck.Videos, as a rich data source, contain knowledge of behaviors, physics, and semantics, but extracting control-specific information from them is challenging due to the lack of action labels.In this work, we introduce a novel framework, Any-point Trajectory Modeling (ATM), that utilizes video demonstrations by pre-training a trajectory model to predict future trajectories of arbitrary points within a video frame.Once trained, these trajectories provide detailed control guidance, enabling the learning of robust visuomotor policies with minimal action-labeled data.Across over 130 language-conditioned tasks we evaluated in both simulation and the real world, ATM outperforms strong video pre-training *First three authors contributed equally: Chuan Wen led the implementation and experiments.Xingyu Lin came up with the idea, supervised the technical development, and contributed to model debugging.John So implemented the Robot-to-robot transfer experiments and UniPi baselines.baselines by 80% on average.Furthermore, we show effective transfer learning of manipulation skills from human videos and videos from a different robot morphology.Visualizations and code are available at: https://xingyu-lin.github.io/atm. I. INTRODUCTIONComputer vision and natural language understanding have made significant advances in recent years [22,7], where the availability of large datasets plays a critical role.Similarly, in robotics, scaling up human demonstration data has been key for learning new skills [6,34,14], with a clear trend of performance improvement with larger datasets [29,6].However, human demonstrations, typically action-labeled trajectories collected via teleoperation devices [55,52], are time-consuming and labor-intensive to collect.For instance, collecting 130K trajectories in [6] took 17 months, making data collection a major bottleneck in robot learning.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AISimulation Techniques and Applications
Traffic Prediction and Management Techniques · Bayesian Modeling and Causal Inference
参考文献 0
施引文献 42
按被引量排序,此处列出前 3 条