计算机科学
编码器
人工智能
单眼
卷积神经网络
背景(考古学)
人工神经网络
模式识别(心理学)
计算复杂性理论
机器学习
计算机视觉
数据挖掘
算法
生物
操作系统
古生物学
作者
Dongdong Zhang,Chunping Wang,Qiang Fu
标识
DOI:10.1117/1.jei.33.6.063045
摘要
We address the problem of self-supervised depth estimation. Given the significance of long-range correlation in-depth estimation, we propose to use the Segment Anything Model (SAM) encoder, a special Vision Transformer (ViT), to model the global context for accurate depth estimation. We also employ a convolutional neural network encoder to assist the network in gathering local information as ViT lacks spatial inductive bias in modeling local information. However, independent encoders lead to insufficient aggregation among features. To compensate for this deficiency, we design a heterogeneous fusion module that facilitates feature fusion by modeling the affinity among heterogeneous features. Due to the unbearable computational burden imposed by the introduction of SAM, we substitute the original SAM with a lightweight variant of SAM to reduce the complexity of the entire network. Extensive experiments on the KITTI dataset show that our proposed model achieves the most competitive results.
科研通智能强力驱动
Strongly Powered by AbleSci AI