计算机科学
机器翻译
自然语言处理
人工智能
翻译(生物学)
语音识别
编码(集合论)
程序设计语言
生物
生物化学
基因
信使核糖核酸
集合(抽象数据类型)
作者
Ramakrishna Appicharla,Kamal Gupta,Asif Ekbal,Pushpak Bhattacharyya
摘要
ABSTRACT This paper studies neural machine translation (NMT) of code‐mixed (CM) text. Specifically, we generate synthetic CM data and how it can be used to improve the translation performance of NMT through the data augmentation strategy. We conduct experiments on three data augmentation approaches viz. CM‐Augmentation, CM‐Concatenation, and Multi‐Encoder approaches, and the latter two approaches are inspired by document‐level NMT, where we use synthetic CM data as context to improve the performance of the NMT models. We conduct experiments on three language pairs, viz. Hindi–English, Telugu–English and Czech–English. Experimental results demonstrate that the proposed approaches significantly improve performance over the baseline model trained without data augmentation and over the existing data augmentation strategies. The CM‐Concatenation model attains the best performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI