遥感
计算机科学
适配器(计算)
分割
人工智能
计算机视觉
遥感应用
图像分割
高光谱成像
语义映射
上下文图像分类
图像处理
作者
Ziyue Qi,Qirui Guo,Libo Wang,Yicong Zhou
标识
DOI:10.1109/tgrs.2026.3702194
摘要
Multimodal remote sensing image semantic segmentation has benefited remarkably from recent progress in deep learning. However, most existing methods treat heterogeneous modalities with excessive uniformity, failing to fully exploit their complementary features. Furthermore, the absence of visual prior knowledge limits their generalization capabilities across diverse scenarios. To address these challenges, we propose a novel multimodal fine-tuning framework based on vision foundation models, namely MultiMoE. Specifically, we select the vision foundation model (DINOv3) as the frozen encoder and develop lightweight Mixture-of-Experts (MoE) adapters on the encoder to transfer general visual knowledge to task-specific semantic contexts. These MoE-based adapters dynamically route samples to dedicated expert pools, thereby processing heterogeneous modalities distinctly and promoting complementary feature representation. Extensive experiments on two widely-used benchmark datasets (ISPRS Vaihingen and Potsdam) demonstrate the competitive performance of our MultiMoE compared to state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI