已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

Differentiating translational English from original English using large language models

计算机科学 自然语言处理 人工智能 变压器 语言模型 编码器 集合(抽象数据类型) 英语 语言学 字错误率 误差分析 语料库 源文本 机器翻译 语音识别 第二语言 自然语言 I类和II类错误 平行语料库 语料库语言学
作者
Ruoyu Hu,Gui Wang,Bin Shao
出处
期刊:Across Languages and Cultures [Akadémiai Kiadó]
卷期号:27 (1): 57-82
标识
DOI:10.1556/084.2025.01149
摘要

Abstract Translated language often carries subtle linguistic markers that set it apart from text originally written in the target language. This study investigates the potential of powerful AI language models to automatically identify these differences, specifically focusing on distinguishing translated texts from original texts. More specifically, this study utilized transformer-based large language models to differentiate translational English and original English. FLOB (Freiburg–LOB Corpus of British English) was selected as the original English corpus, and COCE (Corpus of Chinese-English) was selected as the translational English corpus. Two models were tested: Bidirectional Encoder Representations from Transformers (BERT) and Robustly Optimized BERT Pretraining Approach (RoBERTa). The factors affecting classification results are investigated through SHAP analysis and analysis of text types that have significantly different error rates. Results show that both models achieved excellent performance, with F1-scores of .864 for BERT and .998 for RoBERTa. The text types miscellaneous, general fiction, skills trade, and hobbies, and humor exhibit significantly higher error rates. Reportage, review, and science exhibit significantly lower error rates. Through SHAP analysis, we find that higher error rates may be attributed to simpler sentences in these three text types and the shared characteristics of translational texts, such as the tendencies for simplification and explicitation. Conversely, lower error rates were associated with text types that did not share these characteristics. In summary, transformer-based large language models show great potential for the automatic classification and analysis of translational sentences.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
万花谷发布了新的文献求助10
2秒前
3秒前
4秒前
江子川发布了新的文献求助10
4秒前
大个应助chenchen采纳,获得10
5秒前
6秒前
9秒前
哈哈哈发布了新的文献求助10
9秒前
Orange应助小白采纳,获得20
10秒前
领导范儿应助科研通管家采纳,获得10
10秒前
现代的擎苍完成签到,获得积分0
11秒前
HeJiangle发布了新的文献求助10
12秒前
海阔天空完成签到 ,获得积分10
14秒前
天天发布了新的文献求助10
15秒前
16秒前
Monster完成签到,获得积分10
16秒前
MedChai完成签到,获得积分10
16秒前
16秒前
xzh完成签到,获得积分10
18秒前
拿铁小笼包完成签到,获得积分10
20秒前
温暖的忆霜完成签到,获得积分10
20秒前
共享精神应助湘崽丫采纳,获得10
20秒前
默默樱桃发布了新的文献求助10
22秒前
隐形曼青应助故园无此声采纳,获得10
23秒前
研友_VZG7GZ应助ypyue采纳,获得10
24秒前
孙晓婷完成签到,获得积分10
24秒前
秋天完成签到,获得积分20
24秒前
Mu完成签到,获得积分10
25秒前
26秒前
hui完成签到 ,获得积分10
28秒前
爆米花应助天天采纳,获得10
28秒前
默默樱桃完成签到,获得积分10
30秒前
独特从蓉发布了新的文献求助10
31秒前
32秒前
彭于晏应助pp采纳,获得10
32秒前
小白完成签到,获得积分10
33秒前
搜集达人应助ayin2333采纳,获得10
33秒前
领导范儿应助sweet采纳,获得10
34秒前
34秒前
高分求助中
Markov Chain Monte Carlo 10000
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Common Foundations of American and East Asian Modernisation: From Alexander Hamilton to Junichero Koizumi 5000
Pediatric Dermoscopy Trichoscopy & Onychoscopy 1000
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 700
Matrix Methods in Data Mining and Pattern Recognition Second Edition 610
International Security Studies and Technology :Approaches, Assessments, and Frontiers 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7571390
求助须知:如何正确求助?哪些是违规求助? 9150965
关于积分的说明 19572513
捐赠科研通 7156467
什么是DOI,文献DOI怎么找? 3264048
关于科研通互助平台的介绍 2429357
邀请新用户注册赠送积分活动 2254183