对跨语文转让方法的成本效益分析 (A cost-benefit analysis of cross-lingual transfer methods)

An effective method for cross-lingual transfer is to fine-tune a bilingual or multilingual model on a supervised dataset in one language and evaluating it on another language in a zero-shot manner. Translating examples at training time or inference time are also viable alternatives. However, there are costs associated with these methods that are rarely addressed in the literature. In this work, we analyze cross-lingual methods in terms of their effectiveness (e.g., accuracy), development and deployment costs, as well as their latencies at inference time. Our experiments on three tasks indicate that the best cross-lingual method is highly task-dependent. Finally, by combining zero-shot and translation methods, we achieve the state-of-the-art in two of the three datasets used in this work. Based on these results, we question the need for manually labeled training data in a target language. Code and translated datasets are available at https://github.com/unicamp-dl/cross-lingual-analysis

翻译：一种有效的跨语言传输方法是,对一种语文的受监督数据集的双语或多语文模式进行微调,以零点方式对一种语文进行评价,对另一种语文进行评价;在培训时间或推论时间转换实例也是可行的替代办法;然而,这些方法的费用很少在文献中讨论;在这项工作中,我们分析跨语言方法的有效性(例如准确性)、开发和部署费用及其在推论时间的拖延。我们在三项任务上的实验表明,最佳的跨语言方法高度依赖任务。最后,通过将零点和翻译方法结合起来,我们在这项工作使用的三种数据集中的两种中达到最新水平。根据这些结果,我们质疑是否需要用目标语言进行人工标记的培训数据。代码和翻译数据集可在https://github.com/unamp-dl/crosy-languy-assy-assul查阅。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

“CVPR 2021 接受论文列表 1663篇论文都在这了

专知会员服务

32+阅读 · 2021年6月12日

注意力机制综述

专知会员服务

83+阅读 · 2021年1月26日

【综述：心理学、神经科学和机器学习中的注意力】《Attention in Psychology, Neuroscience, and Machine Learning | Frontiers in Computational Neuroscience》

专知会员服务

41+阅读 · 2020年4月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日