计算机科学
人工智能
变压器
弹丸
计算机视觉
图像分辨率
零(语言学)
遥感
模式识别(心理学)
地质学
物理
电压
量子力学
哲学
化学
语言学
有机化学
作者
Rambabu Damalla,C. Gayathri,Rajeshreddy Datla,Sobhan Babu
标识
DOI:10.1109/lgrs.2025.3554501
摘要
Zero-shot scene classification in remote sensing images presents considerable challenges, primarily due to the diverse variations in scene content and the inconsistent spatial resolutions, which complicate the classification of unseen scene categories. We propose SuperCLIP, a comprehensive framework that integrates a super-resolution module, contrastive language-image pretraining (CLIP), a semantic attribute-guided transformer (SAT), and a visual-semantic projection network (VSPN) to address these challenges. SuperCLIP leverages semantic attributes from three widely used remote sensing scene classification datasets to extract insightful semantic knowledge effectively through CLIP. We use the super-resolution module to obtain high-quality visual scenes of remote sensing images. The SAT enhances the transferability of visual features between seen and unseen categories by localizing object attributes, thereby improving the learning of distinct visual representations. These learned features are further mapped into a semantic embedding space using the VSPN, enabling stronger visual-semantic interactions for more accurate classification. Through extensive experiments, we demonstrate that the SuperCLIP framework significantly improves the classification performance of unseen scene categories across the three benchmark remote sensing datasets, highlighting its effectiveness. The code is available athttps://github.com/ZSL-RSI-SC/SuperCLIP.
科研通智能强力驱动
Strongly Powered by AbleSci AI