计算机科学
人工智能
计算机视觉
上下文图像分类
变压器
模式识别(心理学)
图像(数学)
工程类
电气工程
电压
作者
Ji Zhang,Zhihao Chen,Yiyuan Ge,Mingxin Yu
标识
DOI:10.1109/icicml60161.2023.10424909
摘要
This paper introduces an innovative and efficient multi-scale Vision Transformer (ViT) for the task of image classification. The proposed model leverages the inherent power of transformer architecture and combines it with the concept of multi-scale processing generally used in convolutional neural networks (CNNs). The work aims to address the limitations of conventional ViTs which typically operate at a single scale, hence overlooking the hierarchical structure in visual data. The multi-scale ViT enhances classification performance by processing image features at different scales, effectively capturing both low-level and high-level semantic information. Extensive experimental results demonstrate the superior performance of the proposed model over standard ViTs and other state-of-the-art image classification methods, signifying the effectiveness of the multi-scale approach. This research opens new avenues for incorporating scale-variance in transformer-based models for improved performance in vision tasks.
科研通智能强力驱动
Strongly Powered by AbleSci AI