弹丸
计算机科学
人工智能
认知心理学
心理学
材料科学
冶金
标识
DOI:10.1109/tetci.2025.3592937
摘要
Few-shot learning in image classification aims at learning a model to identify novel categories from a few training samples. Many few-shot learning methods have been proposed to improve the model performance. However, they still suffer from data dependencies, weak feature representation ability, and poor performance on complex tasks. To address these three issues, we propose a global and local attention-based multiscale prototypical network (called GLAP), which includes a feature extraction network and a multiscale classification head. The feature extraction network is stacked by four global attention modules (GAMs) and four local attention modules (LAMs). On the one hand, each GAM learns the global context information in the whole image by calculating the global attention map of all pixels, thereby extracting general features to alleviate data dependencies. On the other hand, each LAM learns the local correspondences between features extracted by the corresponding GAM and features extracted by pre-trained ResNet50 at different scales, which combines high-level semantic features with low-level visual features to enhance feature representation ability. Afterward, the multiscale classification head is designed, which adopts a multiscale metric-based meta-learning (MMML) training method. MMML first calculates the category prototypes of four scales for metric classification, with the aim of generating intermediate predictions of four scales. Then, MMML generates the classification result by combining the intermediate predictions of four scales, thus improving the model performance on complex tasks. Extensive experiments on three benchmark datasets demonstrate that GLAP performs better than state-of-the-art few-shot learning methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI