机器翻译
计算机科学
翻译(生物学)
布鲁
人工智能
钥匙(锁)
适应(眼睛)
自然语言处理
机器学习
度量(数据仓库)
语言模型
极限(数学)
稀缺
主流
机器翻译评价
基础(拓扑)
基于实例的机器翻译
对比度(视觉)
机器翻译软件可用性
语言翻译
还原(数学)
标识
DOI:10.1109/icaice68195.2025.11382383
摘要
Chinese-Tibetan machine translation plays a critical role in preserving Tibetan cultural heritage and promoting cross-lingual bilingual communication, yet it faces three key challenges in practical application: the scarcity of high-quality parallel corpora that limit model training, high computational costs of full-model fine-tuning, and insufficient cross-lingual adaptation capabilities caused by structural differences between Chinese and Tibetan. To resolve these constraints, this study applies the Low-Rank Adaptation (LoRA) lightweight tuning technique on the LlamaFactory platform, selecting two lightweight large language models (LLMs)—Qwen3-1.7B-base and DeepSeek-R1-DistillQwen-1.5B—as base models, and conducting supervised fine-tuning separately for the Tibetan→Chinese and Chinese→Tibetan translation directions; performance evaluation adopts four mainstream machine translation metrics (BLEU, ROUGE-1, ROUGE-2, ChrF) to comprehensively measure translation accuracy and fluency. Experimental results show that LoRA fine-tuning significantly improves the performance of both models: the DeepSeek model’s BLEU score increases from 1.55 to 18.11 in the Tibetan→Chinese direction and from 4.17 to 29.14 in the Chinese→Tibetan direction, while the Qwen model’s BLEU score rises from 0.97 to 3.17 in the Tibetan→Chinese direction and from 2.83 to 14.78 in the Chinese→Tibetan direction; notably, LoRA only tunes 0.0019% of the base model’s parameters (approximately 32k parameters), achieving performance comparable to full-model fine-tuning while reducing computational costs by over 99%, which effectively addresses the cost issue in low-resource language translation model tuning. This study makes three key contributions: it proposes a LoRA-based fine-tuning framework adapted to low-resource Chinese-Tibetan translation scenarios and verifies its effectiveness through experiments, conducts the first comparative study on the translation performance of two lightweight LLMs in bidirectional Chinese-Tibetan translation to provide a reference for subsequent model selection in related fields, and constructs a small-scale yet high-quality bilingual corpus covering news and daily dialogue fields—this corpus enriches data support for low-resource minority language translation research, while the work as a whole demonstrates the application potential of lightweight tuning techniques in this domain and provides a replicable research paradigm for studies on other minority language pairs such as Mongolian-Chinese and Uyghur-Chinese.
科研通智能强力驱动
Strongly Powered by AbleSci AI