G-文件级机器翻译转译员 (G-Transformer for Document-level Machine Translation)

Document-level MT models are still far from satisfactory. Existing work extend translation unit from single sentence to multiple sentences. However, study shows that when we further enlarge the translation unit to a whole document, supervised training of Transformer can fail. In this paper, we find such failure is not caused by overfitting, but by sticking around local minima during training. Our analysis shows that the increased complexity of target-to-source attention is a reason for the failure. As a solution, we propose G-Transformer, introducing locality assumption as an inductive bias into Transformer, reducing the hypothesis space of the attention from target to source. Experiments show that G-Transformer converges faster and more stably than Transformer, achieving new state-of-the-art BLEU scores for both non-pretraining and pre-training settings on three benchmark datasets.

翻译：文件层面的MT模型仍然远远不能令人满意。现有的工作将翻译单位从单句扩大到多个句子。但是,研究表明,当我们进一步将翻译单位扩大到整个文件时,对变异器的监督培训可能失败。在本文中,我们发现这种失败不是由于超装造成的,而是在培训期间停留在本地迷你地带。我们的分析表明,目标对源的注意力日益复杂是失败的原因之一。作为一个解决方案,我们提议G-Transer(Terrafer)将地点假设作为变异器的一种诱导偏差,从而将目标关注的假设空间从目标转向源。实验表明,G- Transer(G-Transer)比变异器更快、更刺切,在三个基准数据集的未准备和训练前环境都达到了新的最先进的BLEU分数。

相关内容

Machine Translation

关注 209

机器翻译（Machine Translation）涵盖计算语言学和语言工程的所有分支，包含多语言方面。特色论文涵盖理论，描述或计算方面的任何下列主题:双语和多语语料库的编写和使用，计算机辅助语言教学，非罗马字符集的计算含义，连接主义翻译方法，对比语言学等。官网地址：http://dblp.uni-trier.de/db/journals/mt/

2021机器学习研究风向是啥？MLP→CNN→Transformer→MLP！

专知会员服务

67+阅读 · 2021年5月23日

最新《Transformers模型》教程，64页ppt

专知会员服务

320+阅读 · 2020年11月26日