高光谱成像
人工智能
计算机科学
计算机视觉
上下文图像分类
变压器
模式识别(心理学)
图像(数学)
工程类
电压
电气工程
摘要
Vision Transformer (ViT) has been thoroughly explored in hyperspectral image classification (HIC). Nevertheless, current ViT-based approaches still acquire discriminative features rather than pattern features, and the majority of them do not fully contemplate the significance of the central pixel in HIC. In this paper, we propose a masked vision Transformer (MViT) for HIC from the perspective of learning pattern features. First and foremost, MViT endeavors for the first time to introduce the masking operations in HIC to learn more robust pattern features instead of distinguishable features, thereby bestowing it with outstanding generalization performance. Secondly, during the training phase of MViT, when conducting random masking operations on the embedded features, we deliberately retain the embedding corresponding to the central pixel to guarantee the effectiveness of the model and emphasize the importance of the central pixel in HIC. Finally, MViT will deactivate the masking operations during the testing phase and utilize all the embedded features to accomplish the classification task, which endows the model with stronger discriminatory ability. Most crucially, MViT is an extremely lightweight model, and by introducing the masking operations during the training phase, its training speed will become unprecedentedly rapid. Experiments conducted on the WHU-Hi dataset demonstrate that MViT can consistently achieve excellent or even the optimal classification results in comparison with the most advanced methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI