自编码
计算机科学
卷积神经网络
人工智能
对偶(语法数字)
模式识别(心理学)
人工神经网络
文学类
艺术
作者
Fujian Zheng,Hong Huang
标识
DOI:10.1109/icras65818.2025.11108872
摘要
High spatial resolution (HSR) image scene classification, which aims to achieve automated recognition of scene categories by analyzing surface information within images, represents a critical research direction in the field of remote sensing intelligent interpretation. However, existing convolutional neural networks (CNNs)-based methods are limited to capturing local detail information, while vision Transformer (ViT)-based approaches fragment large-scale objects by splitting and flattening images, compromising global structural integrity. To address these issues, a dual-branch network combining CNNs and masked autoencoders (MAE), termed DBCMAE, is proposed, aiming to effectively modeling both local and global information in HSR images, thereby achieving scene classification. Specifically, the CNNs branch is first utilized to extract local detail information of multi-scale objects. Then, the MAE branch reconstructs the masked regions based on the visible regions, thereby improving its capability to retain the global structural information of large-scale objects. Finally, the features extracted from both branches are concatenated along the channel dimension and input into a classifier to integrate complementary local and global features. Experimental results demonstrate that DBCMAE achieves competitive classification performance on the AID and NWPU-RESISC45 datasets, with overall classification accuracies reaching 97.8 % and 99.9 %, respectively.
科研通智能强力驱动
Strongly Powered by AbleSci AI