计算机科学
单眼
人工智能
计算机视觉
斜格
成对比较
模棱两可
单目视觉
变压器
视图合成
钥匙(锁)
噪音(视频)
失真(音乐)
面子(社会学概念)
深度知觉
空间分析
深度图
分割
编码器
立体视觉
作者
Xinrui Zeng,Bin Luo,Shuo Zhang,Wei Wang,Jun Liu,Xin Su
出处
期刊:Remote Sensing
[Multidisciplinary Digital Publishing Institute]
日期:2025-10-06
卷期号:17 (19): 3372-3372
被引量:1
摘要
Self-supervised monocular depth estimation from oblique UAV videos is crucial for enabling autonomous navigation and large-scale mapping. However, existing self-supervised monocular depth estimation methods face key challenges in UAV oblique video scenarios: depth discontinuity from geometric distortion under complex viewing angles, and spatial ambiguity in weakly textured regions. These challenges highlight the need for models that combine global reasoning with geometric awareness. Accordingly, we propose RMTDepth, a self-supervised monocular depth estimation framework for UAV imagery. RMTDepth integrates an enhanced Retentive Vision Transformer (RMT) backbone, introducing explicit spatial priors via a Manhattan distance-driven spatial decay matrix for efficient long-range geometric modeling, and embeds a neural window fully-connected CRF (NeW CRFs) module in the decoder to refine depth edges by optimizing pairwise relationships within local windows. To mitigate noise in COLMAP-generated depth for real-world UAV datasets, we constructed a high-fidelity UE4/AirSim simulation environment, which generated a large-scale precise depth dataset (UAV SIM Dataset) to validate robustness. Comprehensive experiments against seven state-of-the-art methods across UAVID Germany, UAVID China, and UAV SIM datasets demonstrate that our model achieves SOTA performance in most scenarios.
科研通智能强力驱动
Strongly Powered by AbleSci AI