机器翻译
计算机科学
编码器
翻译(生物学)
人工智能
布鲁
代表(政治)
背景(考古学)
语音识别
自然语言处理
计算机视觉
操作系统
法学
化学
古生物学
信使核糖核酸
基因
政治
生物
生物化学
政治学
作者
Turghun Tayir,Lin Li,Bei Li,Jianquan Liu,Kong Aik Lee
标识
DOI:10.1109/tai.2024.3354668
摘要
The main purpose of multimodal machine translation is to improve the quality of translation results by taking the corresponding visual context as an additional input. Recently many studies in neural machine translation have attempted to obtain high-quality multimodal representation of encoder or decoder via attention mechanism. However, attention mechanism does not always accurately identify the decisive input for each prediction, which leads to an unsatisfactory multimodal information fusion. To this end, we propose an encoder-decoder calibration method which can automatically calibrate the image and text fusion representation in the encoder, and find the decisive input to the translation in the decoder. We validate our model on the multimodal machine translation dataset Multi30K. Experimental results show that our method significantly outperforms several recent baselines for both English–German and English–French translation tasks in terms of BLEU and METEOR.
科研通智能强力驱动
Strongly Powered by AbleSci AI