人工智能
RGB颜色模型
计算机视觉
姿势
计算机科学
保险丝(电气)
能见度
三维姿态估计
特征(语言学)
融合
特征提取
模式识别(心理学)
对象(语法)
传感器融合
纹理(宇宙学)
目标检测
任务(项目管理)
可视化
融合机制
图像融合
深度知觉
图像纹理
实体造型
利用
噪音(视频)
代表(政治)
作者
Jiaming Zhou,Haoran Tan,Yaonan Wang,Chaoxu Mu,Qing Zhu,Ajmal Mian
标识
DOI:10.1109/tie.2025.3645469
摘要
Estimating the 6-D pose of transparent objects remains a challenging task due to weak visual cues and incomplete depth information caused by refractions and additional reflections from the backsides of transparent objects. These challenges often lead to inaccurate perception and localization. Existing approaches struggle to jointly enhance RGB and depth features or fully exploit depth information. In this article, we propose a depth-directed method that integrates stable diffusion and Mamba-based RGB-D augmentation to address these limitations. Specifically, we use a stable diffusion model to enhance the visibility of transparent objects in RGB images by recovering texture details and reducing background interference. Then, a Mamba-based network completes the sparse and noisy raw depth maps using guidance from the refined RGB features. To effectively fuse the two modalities, we introduce a multiscale RGB-D fusion strategy, where the completed depth not only provides geometric information but also guides RGB feature extraction. This joint representation leads to more accurate and robust 6-D pose estimation. Experimental results on challenging datasets demonstrate that our method significantly enhances object visibility and improves pose estimation accuracy in complex real-world scenes.
科研通智能强力驱动
Strongly Powered by AbleSci AI