有条件的适应性多任务学习:利用更少参数和更少数据,改进国家劳工规划局的转让学习 (Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less Data)

Multi-Task Learning (MTL) networks have emerged as a promising method for transferring learned knowledge across different tasks. However, MTL must deal with challenges such as: overfitting to low resource tasks, catastrophic forgetting, and negative task transfer, or learning interference. Often, in Natural Language Processing (NLP), a separate model per task is needed to obtain the best performance. However, many fine-tuning approaches are both parameter inefficient, i.e., potentially involving one new model per task, and highly susceptible to losing knowledge acquired during pretraining. We propose a novel Transformer architecture consisting of a new conditional attention mechanism as well as a set of task-conditioned modules that facilitate weight sharing. Through this construction (a hypernetwork adapter), we achieve more efficient parameter sharing and mitigate forgetting by keeping half of the weights of a pretrained model fixed. We also use a new multi-task data sampling strategy to mitigate the negative effects of data imbalance across tasks. Using this approach, we are able to surpass single task fine-tuning methods while being parameter and data efficient (using around 66% of the data for weight updates). Compared to other BERT Large methods on GLUE, our 8-task model surpasses other Adapter methods by 2.8% and our 24-task model outperforms by 0.7-1.0% models that use MTL and single task fine-tuning. We show that a larger variant of our single multi-task model approach performs competitively across 26 NLP tasks and yields state-of-the-art results on a number of test and development sets. Our code is publicly available at https://github.com/CAMTL/CA-MTL.

翻译：多任务学习(MTL)网络已成为在不同任务中传授知识的一个很有希望的方法。然而,多任务学习(MTL)网络必须应对挑战,例如:过度适应低资源任务、灾难性的遗忘、负任务转移或学习干扰。通常,在自然语言处理(NLP)中,每个任务需要一个单独的模型才能获得最佳业绩。然而,许多微调方法既低参数,即每个任务可能涉及一个新模式,而且极易丢失在培训前阶段获得的知识。我们提议了一个全新的竞争性变换器结构,包括一个新的有条件关注机制以及一套便利权重共享的任务调整模块。通过这一构建(超网络调整器),我们实现更有效的参数共享和减轻对每个任务的不同模型的忘却。我们还使用新的多任务数据取样战略来减轻数据不平衡的消极影响。使用这个方法,我们可以超越单一任务调整方法,同时进行参数和数据效率(使用大约66%的数据进行权重更新 )。我们通过SARBBGLMTL的更大任务,在GLMTFS上使用我们S-TRA的模型和S-BBBBBBBBBBBL 测试其他方法,在GLMTUL中,我们S-BS-BS-BRBRBS-BSBS-BSBS-BS-BS-BS-BS-BS-BS-BS-BS-BS-BSBS-BS-BS-BS-R-BS-R-R-R-R-BS-L-L-BTL-BTL-S-S-S-L-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-S-L-S-SDL-L-L-L-SDL-L-L-S-S-S-S-S-L-L-L-L-S-S-S-S-S-S-S-S-S-S-S-S-L-S-S-S-S-S-L-S-S-S-S-S

相关内容

多任务学习

关注 161

多任务学习（MTL）是机器学习的一个子领域，可以同时解决多个学习任务，同时利用各个任务之间的共性和差异。与单独训练模型相比，这可以提高特定任务模型的学习效率和预测准确性。多任务学习是归纳传递的一种方法，它通过将相关任务的训练信号中包含的域信息用作归纳偏差来提高泛化能力。通过使用共享表示形式并行学习任务来实现,每个任务所学的知识可以帮助更好地学习其它任务。

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日