情态动词
计算机科学
单眼
人工智能
计算机视觉
感知
估计
模式识别(心理学)
工程类
材料科学
心理学
系统工程
神经科学
高分子化学
作者
Xiaogang Song,Yuping Tan,Jingyu Ning,Xiaofeng Lu,Xinhong Hei
标识
DOI:10.1109/jsen.2025.3567354
摘要
Unsupervised monocular depth estimation has become a mainstream approach in depth estimation due to its ability to eliminate the reliance on expensive ground truth depth annotations, making it more scalable for real-world applications. However, despite significant progress, two major challenges remain: the difficulty in effectively capturing both global context and fine-grained local details, and the limited generalization ability in complex, dynamic scenes.To address these issues, we propose MIPDepth, an unsupervised monocular depth estimation framework based on multi-modal interaction perception. Our method leverages the complementary strengths of convolutional neural networks (CNNs) and vision transformers (ViTs) to overcome their individual limitations. Specifically, the Multi-Scale Pyramid Interaction Perception Layer (MPIPL) is designed to extract rich multi-scale features, while the Multi-Modal Interaction Perception Module (MIPM) enables effective interaction between CNN and ViT streams, enhancing both local detail preservation and global contextual understanding.Extensive experiments on KITTI and Cityscapes benchmarks demonstrate that MIPDepth achieves state-of-the-art performance, especially in dynamic environments, validating its superiority in accuracy, robustness, and generalization.
科研通智能强力驱动
Strongly Powered by AbleSci AI