Object-centric diffusion policies for real-world robotic-arm imitation learning
Prashant Reddy Kasu, Dugan Um
Texas A&M University – Corpus Christi
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Imitation learning in complex, unstructured environments remains challenging due to the difficulty of grounding perception in physically meaningful representations and the need to model multimodal action distributions. Existing approaches often rely on unstructured pixel-level feature encodings or stochastic latent-variable decoders, which can lead to brittle attention in cluttered scenes. In this work, we present a novel integration of detector-based visual representations with conditional diffusion modeling (DINO + CDP) for real-world robotic imitation learning. Our framework utilizes a DINO object detection transformer to extract spatially-grounded object-query embeddings that serve as the conditioning signal for a diffusion-based policy. A primary contribution of this work is the systematic quantification of how scene complexity-measured via image entropy-affects robotic policy performance. By comparing rigid-object baselines with complex biological plant scenes, we demonstrate that organic morphology induces a measurable increase in pixel-level uncertainty that degrades standard pixel-centric models. Our results show that DINO + CDP mitigates this degradation by grounding action generation in stable object-level features. We evaluate our approach using a fully real-world manipulation dataset collected without simulation or synthetic pre-training. To isolate the impact of our architectural choices, we conduct a comparative study within a unified framework against convolutional (CNN-MLP), transformer-patch (ViT), and latent-variable (DETR + CVAE) variants. Experimental results in a robotic-arm biocell setup demonstrate that object-query-conditioned diffusion significantly improves task success rates, produces smoother trajectories, and exhibits superior robustness to high-entropy visual inputs, establishing a scalable pathway for imitation learning in challenging agricultural domains.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIDomain Adaptation and Few-Shot Learning
Reinforcement Learning in Robotics · Robot Manipulation and Learning
参考文献 20
此处列出前 3 条