计算机科学
计算机视觉
图像分割
人工智能
分割
遥感
变压器
地质学
工程类
电压
电气工程
作者
Kejun Liu,Xuesong Yan,Yuanyuan Liu,Chang Tang,Yibing Zhan,Wei Luo,Wujie Zhou,Hongyan Zhang
标识
DOI:10.1109/tgrs.2025.3586620
摘要
Multimodal remote sensing image segmentation (MRSIS) is important for intelligent remote sensing image (RS) interpretation, which encompasses three distinct tasks: semantic segmentation, instance segmentation, and panoptic segmentation. Existing methods typically address individual tasks with specialized models, limiting generalization and real-world applicability. Multi-task learning approaches have introduced separated task heads to unify tasks, yet we identify two key challenges when directly applying them to MRSIS: (1) the modality gap, arising from semantic discrepancies and granularity discrepancies across RS modalities, and (2) the task gap, due to varying preferences in learning different segmentation tasks. To overcome these challenges, we propose PMTSeg—a novel Prompt-driven Multimodal Transformer for task-adapted MRSIS. PMTSeg integrates three key components: (1) Task-common Multimodal Affinity Approximation (TMAA), (2) Task-common Multi-scale Semantic Fusion (TMSF), and (3) a unified Prompt-driven Segmentation Head (PSH). First, TMAA addresses the modality gap by approximating inter-modal affinity matrices, extracting task-common features across modalities and aligning semantic information. Then, TMSF further integrates these features using the scale-matched fusion at multiple scales to produce enriched, multi-scale task-common features. Moreover, to address the task gap, the PSH leverages task-adapted text prompts and task-adapted contrastive loss to model relationships across tasks, enabling adaptive optimization for robust and universal MRSIS performance. Extensive experiments on three MRSIS datasets—VALID, SEMCITY TOULOUSE, and UBCV2—demonstrate that PMTSeg significantly surpasses state-of-the-art methods in all three segmentation tasks, offering a unified and accurate solution to MRSIS.
科研通智能强力驱动
Strongly Powered by AbleSci AI