StereoAnything: Advanced Zero-Shot Stereo Imaging for Robotic Grasp Detection With Transparent Objects
受人类感知启发,提出一种无需分割先验的框架,利用基础模型特征隐式学习透明物体重建策略,在真实机器人实验中达到96%抓取成功率,且不遗忘不透明物体抓取能力。
Grasping transparent objects remains challenging for robotic systems due to their reflective and refractive properties, which distort depth perception and introduce background noise. Unlike humans, who leverage life experience to perceive depth intuitively, robotic algorithms often fail to generalize across different object types. To address this, we propose a novel framework inspired by human perception for grasping transparent objects. Our approach extends features extracted by foundation models to implicitly learn reconstruction strategies for transparent objects without requiring segmentation priors. Crucially, our framework maintains strong performance across all types of objects and scenes, preventing catastrophic forgetting of opaque objects while learning to perceive transparent ones. By integrating affordance information, our method dynamically guides a five-finger dexterous hand to execute diverse grasping strategies based on human intent. To tackle the challenge of annotating transparent objects, we constructed a large-scale synthetic dataset with depth information, affordance data, and automated annotations. Our framework demonstrates strong generalization, achieving a 96% grasp success rate in real-world robotic experiments and proving its broad applicability across varied environments.