Efficient Six-Degrees of Freedom (6-DoF) Grasp Pose Detection in Cluttered Scenes via Multimodal Fusion and Object-Centric Receptive Fields
Xiaozheng Liu, Kechen Song, ZhongLei Liu, Zehao Xu, Yunhui Yan
Northeastern University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Efficient 6-DoF grasp pose detection is a fundamental and challenging task in robotic manipulation. Existing methods mainly rely on the geometric features of point clouds to synthesise 6-DoF grasp poses, as point clouds contain rich geometric information. However, the rich visual cues contained in images, such as object edges, textures, and other fine-grained details, can also serve as important references. In this paper, we propose a simple yet efficient multimodal 6-DoF grasp pose detection algorithm that also integrates object-centric receptive fields. Specifically, we design an early fusion strategy to integrate point cloud and image information, providing more discriminative multimodal features for grasp pose detection. Furthermore, an object-centric receptive fields module based on grounded segment anything model (Grounded-SAM) is introduced, which provides global references for each point during the grasp parameter decoding stage. Experimental results show that the proposed model achieves competitive performance on the GraspNet-1Billion dataset. Compared with baseline, the AP improves by 7.44, 8.42, and 2.37 in Seen, Similar, and Novel scenes, respectively. To test the model's robustness to image noise, we designed a series of noise experiments, demonstrating the model's strong anti-interference ability. Finally, we deployed the algorithm on a Franka robot and conducted extensive grasping experiments, achieving a high grasp success rate.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Robot Manipulation and Learning
Human Pose and Action Recognition · Hand Gesture Recognition Systems
参考文献 32
此处列出前 3 条
引用本文 1
按被引量排序,此处列出前 3 条