Universal Manipulation Interface
去安UP
2024年09月17日 21:29
收录于文集
共4篇

录用情况:

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

RSS 2024

作者:

Cheng Chi∗1,2, Zhenjia Xu∗1,2, Chuer Pan1, Eric Cousineau3, 

Benjamin Burchfiel3, Siyuan Feng3, Russ Tedrake3, Shuran Song1,2 

1Stanford University, 2 Columbia University, 3Toyota Research Insititute

动机:

UMI 是一个数据收集和策略学习框架,允许将野外人类演示的技能直接转移到可部署的机器人策略。

“与自然语言处理(NLP)或计算机视觉(CV)等其他领域不同,互联网上没有广泛可用的机器人数据,因此我们必须自己收集数据。”

现有方法及其不足

  1. Collecting targeted in-the-lab robot datasets via teleoperation. (costs for hardware and expert operators)

  2. Leveraging unstructured in-the-wild human videos. (human videos exhibit a large embodiment gap to robots)

  3. Using sensorized hand-held grippers as a data collection interface.(struggle to balance action diversity with transferability)

本文方法属于第三类,What prevents action transfer?

  1. Insufficient visual context: 相机视角受限,尤其是接近被操纵物体的时候。

  2. Action imprecision: SfM相机尺度模糊、运动模糊或纹理不足,极大地限制了系统可以采用的任务的精度。

  3. Latency discrepancies: 推理的延迟会导致策略遇到分布外的输入。

  4. Insufficient policy representation: 简单的策略表示(如回归)难以学到多模态动作分布。

本文的方法

  • 使用鱼眼镜头来增加视场和视觉上下文,添加侧镜来提供隐式立体观测。

  • 结合GoPro内置的IMU传感器,在快速运动下实现鲁棒跟踪。

  • 采用推理时间延迟匹配来处理不同的传感器观察和执行延迟,使用相对轨迹作为动作表示来消除对精确全局动作的需求。

  • 应用Diffusion Policy对多模态动作分布建模。

B站上传图片比较麻烦,全文请移步:

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots - 弃安的文章

https://zhuanlan.zhihu.com/p/720537629