Softmax函数
概化理论
变压器
计算机科学
自然语言处理
过度拟合
语言模型
人工智能
机器学习
语音识别
数学
人工神经网络
统计
物理
电压
量子力学
作者
Shuoran Jiang,Youcheng Pan,Qingcai Chen,Yang Xiang,Xiangping Wu
标识
DOI:10.1109/taslp.2024.3394774
摘要
Although the pre-trained Transformers learned general linguistic knowledge from large-scale corpus, they still over-fit on the lexical biases when fine-tuning on specific datasets. This problem limits the generalizability of pre-trained models, particularly when learning over out-of-distribution (OOD) data. To address this issue, this paper proposes a self-adaptive language masking (AdaLMask) paradigm to fine-tune the pre-trained Transformers. AdaLMask obviates lexical biases by eliminating the dependence on semantically inessential words. Specifically, AdaLMask learns a Gumbel-Softmax distribution to determine the desired masking positions, and the distribution parameters are optimized via a representation-invariant (RInv) objective to ensure the masked positions are semantically lossless. Four natural language processing tasks are chosen to evaluate the effectiveness of the proposed method on the robustness of lexical biases and OOD generalization. All empirical results demonstrate that the AdaLMask paradigm substantially improves the OOD generalization of pre-trained Transformers.
科研通智能强力驱动
Strongly Powered by AbleSci AI