人工智能
计算机科学
单眼
计算机视觉
点云
编码器
编码(集合论)
不变(物理)
模式识别(心理学)
数学
集合(抽象数据类型)
数学物理
程序设计语言
操作系统
作者
Wei Yin,Jianming Zhang,Oliver Wang,Simon Niklaus,Long Mai,Simon Chen,Chunhua Shen
出处
期刊:
日期:2021-06-01
被引量:152
标识
DOI:10.1109/cvpr46437.2021.00027
摘要
Despite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in mixed-data depth prediction training, and possible unknown camera focal length. We investigate this problem in detail, and propose a two-stage framework that first predicts depth up to an unknown scale and shift from a single monocular image, and then use 3D point cloud encoders to predict the missing depth shift and focal length that allow us to recover a realistic 3D scene shape. In addition, we propose an image-level normalized regression loss and a normal-based geometry loss to enhance depth prediction models trained on mixed datasets. We test our depth model on nine unseen datasets and achieve state-of-the-art performance on zero-shot dataset generalization. Code is available at: https://git.io/Depth
科研通智能强力驱动
Strongly Powered by AbleSci AI