CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications

变压器 计算机科学 人工智能 电气工程 工程类 电压
作者
Tianfang Zhang,Lei Li,Yang Zhou,Wentao Liu,Chen Qian,Xiangyang Ji,Ji, Xiangyang
出处
期刊:Cornell University - arXiv [Cornell University]
被引量:17
标识
DOI:10.48550/arxiv.2408.03703
摘要

Vision Transformers (ViTs) mark a revolutionary advance in neural networks with their token mixer's powerful global context capability. However, the pairwise token affinity and complex matrix operations limit its deployment on resource-constrained scenarios and real-time applications, such as mobile devices, although considerable efforts have been made in previous works. In this paper, we introduce CAS-ViT: Convolutional Additive Self-attention Vision Transformers, to achieve a balance between efficiency and performance in mobile applications. Firstly, we argue that the capability of token mixers to obtain global contextual information hinges on multiple information interactions, such as spatial and channel domains. Subsequently, we propose Convolutional Additive Token Mixer (CATM) employing underlying spatial and channel attention as novel interaction forms. This module eliminates troublesome complex operations such as matrix multiplication and Softmax. We introduce Convolutional Additive Self-attention(CAS) block hybrid architecture and utilize CATM for each block. And further, we build a family of lightweight networks, which can be easily extended to various downstream tasks. Finally, we evaluate CAS-ViT across a variety of vision tasks, including image classification, object detection, instance segmentation, and semantic segmentation. Our M and T model achieves 83.0\%/84.1\% top-1 with only 12M/21M parameters on ImageNet-1K. Meanwhile, throughput evaluations on GPUs, ONNX, and iPhones also demonstrate superior results compared to other state-of-the-art backbones. Extensive experiments demonstrate that our approach achieves a better balance of performance, efficient inference and easy-to-deploy. Our code and model are available at: \url{https://github.com/Tianfang-Zhang/CAS-ViT}
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
aaaa应助橄榄油采纳,获得20
1秒前
Dawang发布了新的文献求助10
1秒前
mudong关注了科研通微信公众号
2秒前
文献期待应助xiaokezhang采纳,获得10
4秒前
余白薇完成签到,获得积分20
5秒前
5秒前
顾矜应助msd2phd采纳,获得10
6秒前
药膳干完成签到,获得积分10
6秒前
6秒前
忧伤的含海完成签到,获得积分10
7秒前
隐形曼青应助Yingzi采纳,获得10
8秒前
8秒前
子勿语完成签到 ,获得积分10
9秒前
lw完成签到,获得积分10
9秒前
杨柳完成签到,获得积分10
10秒前
gao应助luo采纳,获得10
11秒前
12秒前
陈丽发布了新的文献求助10
12秒前
李爱国应助花南星采纳,获得10
12秒前
余白薇发布了新的文献求助10
13秒前
HCody完成签到 ,获得积分10
14秒前
14秒前
XinG完成签到,获得积分10
17秒前
Yingzi完成签到,获得积分10
18秒前
Lucas应助venti采纳,获得10
19秒前
19秒前
欣观完成签到,获得积分10
19秒前
小恩发布了新的文献求助10
19秒前
XY发布了新的文献求助10
20秒前
23秒前
25秒前
科研通AI6.2应助lm18994782585采纳,获得10
26秒前
花南星发布了新的文献求助10
26秒前
秋风举报yn求助涉嫌违规
27秒前
Dawang完成签到,获得积分10
28秒前
28秒前
陈皮完成签到 ,获得积分10
29秒前
欣观发布了新的文献求助10
29秒前
Fish发布了新的文献求助10
30秒前
睡睡林关注了科研通微信公众号
30秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Encyclopedia of Cardiovascular Research and Medicine(2e) 820
自動車の空力技術 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7781791
求助须知:如何正确求助?哪些是违规求助? 9321417
关于积分的说明 20382975
捐赠科研通 7369678
什么是DOI,文献DOI怎么找? 3320126
关于科研通互助平台的介绍 2467955
邀请新用户注册赠送积分活动 2336049