计算机科学
人工智能
平滑的
任务(项目管理)
弹丸
自然语言理解
自然语言处理
训练集
一般化
机器学习
班级(哲学)
水准点(测量)
正规化(语言学)
理论(学习稳定性)
自然语言
数学
工程类
有机化学
地理
系统工程
化学
数学分析
计算机视觉
大地测量学
作者
Meng Yu,Jiaxin Huang,Yu Zhang,Jiawei Han
标识
DOI:10.48550/arxiv.2202.04538
摘要
Pretrained language models (PLMs) have demonstrated remarkable performance in various natural language processing tasks: Unidirectional PLMs (e.g., GPT) are well known for their superior text generation capabilities; bidirectional PLMs (e.g., BERT) have been the prominent choice for natural language understanding (NLU) tasks. While both types of models have achieved promising few-shot learning performance, their potential for zero-shot learning has been underexplored. In this paper, we present a simple approach that uses both types of PLMs for fully zero-shot learning of NLU tasks without requiring any task-specific data: A unidirectional PLM generates class-conditioned texts guided by prompts, which are used as the training data for fine-tuning a bidirectional PLM. With quality training data selected based on the generation probability and regularization techniques (label smoothing and temporal ensembling) applied to the fine-tuning stage for better generalization and stability, our approach demonstrates strong performance across seven classification tasks of the GLUE benchmark (e.g., 72.3/73.8 on MNLI-m/mm and 92.8 on SST-2), significantly outperforming zero-shot prompting methods and achieving even comparable results to strong few-shot approaches using 32 training samples per class.
科研通智能强力驱动
Strongly Powered by AbleSci AI