遥感
计算机科学
特征(语言学)
人工智能
模态(人机交互)
对偶(语法数字)
扩散
模式识别(心理学)
特征提取
遥感应用
蒸馏
上下文图像分类
计算机视觉
作者
Dan Xu,Wenqian Dong,Song Xiao,Jiahui Qu,Y F Li
标识
DOI:10.1109/tgrs.2026.3697497
摘要
In recent years, research on joint classification of multi-modal remote sensing data has achieved remarkable progress. However, constrained by imaging conditions or sensor issues, missing modality frequently occurs in practice. Knowledge distillation, as one of the mainstream strategies to deal with missing modality, is widely used due to its excellent transfer ability. However, most existing knowledge distillation methods focus on directly transferring available modal information through feature alignment or predictive constraints, without explicitly modeling or refining the latent semantic structure of missing modality. Furthermore, the heterogeneity between different modalities makes effective alignment difficult through direct transfer, thus limiting the utilization of cross-modal complementary information and consequently hindering model performance improvement in remote sensing scenarios with missing modality. Therefore, in this paper, we propose a diffusion feature completion-driven textual dual distillation model (DFC-TD2) for multi-modal remote sensing image classification with missing modality. This model explicitly models the missing modality, introduces textual semantic information to guide cross-modal feature transfer, and achieves multi-constraint collaborative optimization under the supervision of classification labels, thereby effectively improving the overall performance of multi-modal classification tasks with missing modality. Specifically, a shared-constrained diffusion feature completion network (SCFC) is designed, which uses shared features extracted and frozen from multi-modal data as constraints to guide the diffusion model to effectively model and complete missing modal features, thereby generating semantically consistent and more discriminative modal representations. Furthermore, a text-guided dual distillation network (TGD2) is designed to guide cross-modal feature alignment using text semantics and combine it with classification label supervision to construct a feature-label dual constraint, which effectively improves feature discriminativeness and supervision efficiency, thereby mitigating cross-modal representation heterogeneity and improving the stability and effectiveness of knowledge transfer. Experimental results on three public datasets demonstrate the effectiveness of this method in dealing with missing modality.
科研通智能强力驱动
Strongly Powered by AbleSci AI