计算机科学
机器翻译
平行语料库
自然语言处理
人工智能
语料库
领域(数学分析)
语料库语言学
翻译(生物学)
基于实例的机器翻译
资源(消歧)
语言翻译
情报检索
生物化学
数学
计算机网络
化学
信使核糖核酸
基因
数学分析
作者
Tian Liang,Derek F. Wong,Lidia S. Chao,Paulo Quaresma,Francisco Oliveira,Lu Yi
出处
期刊:Universidade de Évora - Repositorio Universidade de Évora
日期:2014-05-01
卷期号:: 1837-1842
被引量:92
摘要
Parallel corpus is a valuable resource for cross-language information retrieval and data-driven natural language processing systems, especially for Statistical Machine Translation (SMT). However, most existing parallel corpora to Chinese are subject to in-house use, while others are domain specific and limited in size. To a certain degree, this limits the SMT research. This paper describes the acquisition of a large scale and high quality parallel corpora for English and Chinese. The corpora constructed in this paper contain about 15 million English-Chinese (E-C) parallel sentences, and more than 2 million training data and 5,000 testing sentences are made publicly available. Different from previous work, the corpus is designed to embrace eight different domains. Some of them are further categorized into different topics. The corpus will be released to the research community, which is available at the NLP 2 CT 1 website.
科研通智能强力驱动
Strongly Powered by AbleSci AI