高光谱成像
计算机科学
人工智能
上下文图像分类
图像(数学)
计算机视觉
模式识别(心理学)
遥感
地质学
作者
Wenzhen Wang,Fang Liu,Hongyuan Zhu,Liang Xiao
标识
DOI:10.1109/tgrs.2025.3548607
摘要
Land cover in different scenes generally exhibits scene-invariant category semantic, typically represented and described consistently in a textual modality. Traditional cross-scene classification methods often treat categories as discrete class labels, neglecting their semantic information, or use category names merely as auxiliary textual modalities to enhance the discriminative representations of land cover. However, the cross-scene consistency of category semantic for land cover remains underexplored and underutilized. To address this issue, the text-driven adaptive semantic alignment network (TASA-Net) is proposed in this article for cross-scene hyperspectral image classification (HSIC). TASA-Net employs hand-crafted template prompts for stable category descriptions and vision-guided fine semantic prompts (VG-FSPs) for dynamic scene adaptation. Through a dual-gated adaptive mechanism, TASA-Net optimally weights coarse- and fine-grained semantics in a shared space, ensuring stable yet discriminative semantic representation. Additionally, cross-modal semantic alignment projects visual features into the shared semantic space, while a soft alignment strategy dynamically adjusts category correlations to enhance intraclass consistency and mitigate domain shifts. Ultimately, by leveraging text-driven semantic consistency representation, TASA-Net achieves zero-shot cross-scene transfer for unsupervised classification. Experiments demonstrate superior performance across multiple hyperspectral datasets, validating the critical role of textual modality in enhancing model robustness and cross-scene generalization ability.
科研通智能强力驱动
Strongly Powered by AbleSci AI