人工智能
计算机科学
杠杆(统计)
计算机视觉
RGB颜色模型
利用
点云
先验概率
变压器
情态动词
模式识别(心理学)
深度学习
工程类
化学
贝叶斯概率
计算机安全
电压
高分子化学
电气工程
作者
Linqing Zhao,Yi Wei,Jianqin Li,Jie Zhou,Jiwen Lu
标识
DOI:10.1109/tip.2024.3355807
摘要
In this paper, we present a Structure-aware Cross-Modal Transformer (SCMT) to fully capture the 3D structures hidden in sparse depths for depth completion. Most existing methods learn to predict dense depths by taking depths as an additional channel of RGB images or learning 2D affinities to perform depth propagation. However, they fail to exploit 3D structures implied in the depth channel, thereby losing the informative 3D knowledge that provides important priors to distinguish the foreground and background features. Moreover, since these methods rely on the color textures of 2D images, it is challenging for them to handle poor-texture regions without the guidance of explicit 3D cues. To address this, we disentangle the hierarchical 3D scene-level structure from the RGB-D input and construct a pathway to make sharp depth boundaries and object shape outlines accessible to 2D features. Specifically, we extract 2D and 3D features from depth inputs and the back-projected point clouds respectively by building a two-stream network. To leverage 3D structures, we construct several cross-modal transformers to adaptively propagate multi-scale 3D structural features to the 2D stream, energizing 2D features with priors of object shapes and local geometries. Experimental results show that our SCMT achieves state-of-the-art performance on three popular outdoor (KITTI) and indoor (VOID and NYU) benchmarks.
科研通智能强力驱动
Strongly Powered by AbleSci AI