将神经数学错误校正作为低资源机器翻译任务处理 (Approaching Neural Grammatical Error Correction as a Low-Resource Machine Translation Task)

Previously, neural methods in grammatical error correction (GEC) did not reach state-of-the-art results compared to phrase-based statistical machine translation (SMT) baselines. We demonstrate parallels between neural GEC and low-resource neural MT and successfully adapt several methods from low-resource MT to neural GEC. We further establish guidelines for trustable results in neural GEC and propose a set of model-independent methods for neural GEC that can be easily applied in most GEC settings. Proposed methods include adding source-side noise, domain-adaptation techniques, a GEC-specific training-objective, transfer learning with monolingual data, and ensembling of independently trained GEC models and language models. The combined effects of these methods result in better than state-of-the-art neural GEC models that outperform previously best neural GEC systems by more than 10% M$^2$ on the CoNLL-2014 benchmark and 5.9% on the JFLEG test set. Non-neural state-of-the-art systems are outperformed by more than 2% on the CoNLL-2014 benchmark and by 4% on JFLEG.

翻译：与基于词基的统计机器翻译(SMT)基线相比,古典错误校正(GEC)的神经方法没有达到最先进的结果。我们展示了神经GEC和低资源神经MT之间的平行,并成功地将低资源MT的几种方法与神经GEC相适应。我们进一步为神经GEC的可信任结果制定了指导方针,并提出了一套可在大多数GEC环境中轻易应用的神经GEC模型独立方法。建议的方法包括增加源侧噪音、域适应技术、GEC特定培训目标、用单语数据传输学习以及集成独立培训的GEC模型和语言模型。这些方法的综合效果使得比最先进的神经EC模型的更好,该模型在CONLL-2014基准中比以前最好的神经EC系统高出10%以上,在JFLEG测试中比4.9%高。在CONLU-2014基准中,NLU-NLF基准比4%以上。

相关内容

Machine Translation

关注 209

机器翻译（Machine Translation）涵盖计算语言学和语言工程的所有分支，包含多语言方面。特色论文涵盖理论，描述或计算方面的任何下列主题:双语和多语语料库的编写和使用，计算机辅助语言教学，非罗马字符集的计算含义，连接主义翻译方法，对比语言学等。官网地址：http://dblp.uni-trier.de/db/journals/mt/

多语言神经机器翻译综述论文，34页pdf，A Comprehensive Survey of Multilingual Neural Machine Translation

专知会员服务

19+阅读 · 2020年4月25日

因果图，Causal Graphs，52页ppt

专知会员服务

250+阅读 · 2020年4月19日

【Google】无监督机器翻译，Unsupervised Machine Translation

专知会员服务

36+阅读 · 2020年3月3日

【ICLR2020】理解非自回归机器翻译中的知识蒸馏（Understanding Knowledge Distillation in Non-autoregressive Machine Translation）

专知会员服务

11+阅读 · 2019年12月28日