计算机科学
判别式
一致性(知识库)
人工智能
任务(项目管理)
生成语法
安全性令牌
稳健性(进化)
下游(制造业)
生成模型
机器学习
生物化学
化学
运营管理
计算机安全
管理
经济
基因
作者
Ning Liao,Bowen Shi,Min Cao,Xiaopeng Zhang,Qi Tian,Junchi Yan
标识
DOI:10.48550/arxiv.2303.04998
摘要
Prompt learning has achieved great success in efficiently exploiting large-scale pre-trained models in natural language processing (NLP). It reformulates the downstream tasks as the generative pre-training ones to achieve consistency, thus improving the performance stably. However, when transferring it to the vision area, current visual prompt learning methods are almost designed on discriminative pre-trained models, and there is also a lack of careful design to unify the forms of pre-training and downstream tasks. To explore prompt learning on the generative pre-trained visual model, as well as keeping the task consistency, we propose Visual Prompt learning as masked visual Token Modeling (VPTM) to transform the downstream visual classification into the pre-trained masked visual token prediction. In addition, we develop the prototypical verbalizer for mapping the predicted visual token with implicit semantics to explicit downstream labels. To our best knowledge, VPTM is the first visual prompt method on the generative pre-trained visual model, which achieves consistency between pre-training and downstream visual classification by task reformulation. Experiments show that VPTM outperforms other visual prompt methods and achieves excellent efficiency. Moreover, the task consistency of VPTM contributes to the robustness against prompt location, prompt length and prototype dimension, and could be deployed uniformly.
科研通智能强力驱动
Strongly Powered by AbleSci AI