同样的培训前损失, " 更好的下游:语言模式的隐含偏见问题 " (Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models)

Language modeling on large-scale datasets leads to impressive performance gains on various downstream language tasks. The validation pre-training loss (or perplexity in autoregressive language modeling) is often used as the evaluation metric when developing language models since the pre-training loss tends to be well-correlated with downstream performance (which is itself difficult to evaluate comprehensively). Contrary to this conventional wisdom, this paper shows that 1) pre-training loss cannot fully explain downstream performance and 2) flatness of the model is well-correlated with downstream performance where pre-training loss is not. On simplified datasets, we identify three ways to produce models with the same (statistically optimal) pre-training loss but different downstream performance: continue pre-training after convergence, increasing the model size, and changing the training algorithm. These experiments demonstrate the existence of implicit bias of pre-training algorithms/optimizers -- among models with the same minimal pre-training loss, they implicitly prefer more transferable ones. Toward understanding this implicit bias, we prove that SGD with standard mini-batch noise implicitly prefers flatter minima in language models, and empirically observe a strong correlation between flatness and downstream performance among models with the same minimal pre-training loss. We also prove in a synthetic language setting that among the models with the minimal pre-training loss, the flattest model transfers to downstream tasks.

翻译：在大型数据集上建模的语言模型导致在各种下游语文任务上取得令人印象深刻的业绩收益。在培训前的认证损失(或自动递减语言模型的混乱性)往往被用作开发语言模型的评价衡量标准,因为培训前的损失往往与下游业绩密切相关(这本身难以全面评估)。与传统智慧相反,本文表明:(1) 培训前的损失不能充分解释下游业绩;(2) 培训前的损失与该模式的平整性在培训前损失不明显的情况下与下游业绩密切相关。在简化的数据集方面,我们确定三种方法,以同样的(统计上最理想的)培训前损失和不同的下游业绩制作模型:在培训前继续培训前的整合,增加模式的规模,改变培训算法。这些实验表明,在培训前算法/优化者之间存在着隐含的偏差,在培训前损失最小的模型中,它们隐含着更可转让的模式。为了理解这种隐含的偏差,我们证明,标准微批噪音SGD偏差性在语言模型中隐含自上最优的迷,在培训前损失前,但下游表现为最低的下游模式。在最低的模型中也以经验上证明,在最低损失模型之间确立了。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日