融合语言模型权重的无数据知识融合 (Dataless Knowledge Fusion by Merging Weights of Language Models) - 专知论文

会员服务 ·

0

融合 · 知识融合 · 训练数据 · 知识 · 语言模型 ·

2023 年 4 月 5 日

Dataless Knowledge Fusion by Merging Weights of Language Models

翻译：融合语言模型权重的无数据知识融合

Xisen Jin,Xiang Ren,Daniel Preotiuc-Pietro,Pengxiang Cheng

from arxiv, ICLR 2023; The code is available at https://github.com/bloomberg/dataless-model-merging

Fine-tuning pre-trained language models has become the prevalent paradigm for building downstream NLP models. Oftentimes fine-tuned models are readily available but their training data is not, due to data privacy or intellectual property concerns. This creates a barrier to fusing knowledge across individual models to yield a better single model. In this paper, we study the problem of merging individual models built on different training data sets to obtain a single model that performs well both across all data set domains and can generalize on out-of-domain data. We propose a dataless knowledge fusion method that merges models in their parameter space, guided by weights that minimize prediction differences between the merged model and the individual models. Over a battery of evaluation settings, we show that the proposed method significantly outperforms baselines such as Fisher-weighted averaging or model ensembling. Further, we find that our method is a promising alternative to multi-task learning that can preserve or sometimes improve over the individual models without access to the training data. Finally, model merging is more efficient than training a multi-task model, thus making it applicable to a wider set of scenarios.

翻译：在建立下游NLP模型时，微调预训练语言模型已成为主流范式。但通常情况下，由于数据隐私或知识产权问题，微调模型已经可用，但训练数据不可用。这就为跨越各个模型融合知识以生成更好的单一模型创建了一个障碍。在本文中，我们研究了建立在不同训练数据集上的个别模型的合并问题，以获得跨所有数据集领域都能表现良好的单一模型，并能推广到域外数据。我们提出了一种无数据知识融合方法，该方法在参数空间中合并模型，由权重引导，使合并模型与个别模型之间的预测差异最小化。在一系列评估设置中，我们展示了所提出的方法显著优于基线，例如Fisher加权平均或模型集成。此外，我们发现我们的方法是多任务学习的一个有前途的替代方案，能在没有访问训练数据的情况下保留或有时提高个别模型的表现。最后，模型合并比训练多任务模型更高效，因此适用于更广泛的场景。

0

相关内容

【ICML2022】基于自适应上下文池化的高效表示学习

【ICML2022】基于自适应上下文池化的高效表示学习

专知会员服务

20+阅读 · 2022年7月9日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【知识图谱@EMNLP2020】Knowledge Graphs in NLP @ EMNLP 2020

【知识图谱@EMNLP2020】Knowledge Graphs in NLP @ EMNLP 2020

专知会员服务

43+阅读 · 2020年11月22日

【翻译-ACL2020】使用知识库嵌入改进知识图上的多跳问答

【翻译-ACL2020】使用知识库嵌入改进知识图上的多跳问答

专知会员服务

70+阅读 · 2020年7月3日

【KDD2020】图神经网络生成式预训练，GPT-GNN: Generative Pre-Training of Graph Neural Networks

【KDD2020】图神经网络生成式预训练，GPT-GNN: Generative Pre-Training of Graph Neural Networks

专知会员服务

99+阅读 · 2020年7月3日

【ACL2020】不要停止预训练:根据领域和任务自适应调整语言模型，Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

【ACL2020】不要停止预训练:根据领域和任务自适应调整语言模型，Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

专知会员服务

46+阅读 · 2020年4月25日

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

专知会员服务

96+阅读 · 2020年3月25日

【芝加哥大学】GRAPH-BERT: Only Attention is Needed for Learning Graph Representations

【芝加哥大学】GRAPH-BERT: Only Attention is Needed for Learning Graph Representations

专知会员服务

85+阅读 · 2020年1月15日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

【推荐】用Tensorflow理解LSTM

【推荐】用Tensorflow理解LSTM

机器学习研究会

36+阅读 · 2017年9月11日

孤独症的iPSC模型研究

国家自然科学基金

1+阅读 · 2015年12月31日

儿童孤独症的神经环路机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

重组腺病毒介导Raf基因持续激活Erk1/2/Merk信号通路修复脊髓损伤的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

大白菜KIN基因的表达及其pre-mRNA加工机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

CPU/GPGPU紧耦合异构多核系统共享Last Level Cache优化研究

国家自然科学基金

0+阅读 · 2012年12月31日

金属晶粒长大动力学的多尺度模拟

国家自然科学基金

0+阅读 · 2012年12月31日

SPAC系统中农作物水循环知识融合模型研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于Linked Open Data的Web服务语义互操作关键技术

国家自然科学基金

0+阅读 · 2012年12月31日

慢病毒介导ChABC基因修饰神经干细胞治疗脊髓损伤的研究

国家自然科学基金

0+阅读 · 2011年12月31日

基因修饰的内皮祖细胞靶向治疗HER-2阳性肿瘤的实验研究

国家自然科学基金

0+阅读 · 2009年12月31日

Meta-Learning Online Adaptation of Language Models

Arxiv

0+阅读 · 2023年5月24日

Large Language Models are Better Reasoners with Self-Verification

Arxiv

0+阅读 · 2023年5月23日

To Copy Rather Than Memorize: A Vertical Learning Paradigm for Knowledge Graph Completion

Arxiv

0+阅读 · 2023年5月23日

Inspecting and Editing Knowledge Representations in Language Models

Arxiv

0+阅读 · 2023年5月22日

Text-to-SQL Error Correction with Language Models of Code

Arxiv

0+阅读 · 2023年5月22日

The CLIP Model is Secretly an Image-to-Prompt Converter

Arxiv

0+阅读 · 2023年5月22日

Joint Foundation Model Caching and Inference of Generative AI Services for Edge Intelligence

Arxiv

0+阅读 · 2023年5月20日

Application of Knowledge Distillation to Multi-task Speech Representation Learning

Arxiv

0+阅读 · 2023年5月19日

A One-Class Classifier for the Detection of GAN Manipulated Multi-Spectral Satellite Images

Arxiv

0+阅读 · 2023年5月19日

Which Knowledge Graph Is Best for Me?

Arxiv

11+阅读 · 2018年9月28日

VIP会员

文章信息

相关主题

相关VIP内容

【ICML2022】基于自适应上下文池化的高效表示学习

【ICML2022】基于自适应上下文池化的高效表示学习

专知会员服务

20+阅读 · 2022年7月9日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【知识图谱@EMNLP2020】Knowledge Graphs in NLP @ EMNLP 2020

【知识图谱@EMNLP2020】Knowledge Graphs in NLP @ EMNLP 2020

专知会员服务

43+阅读 · 2020年11月22日

【翻译-ACL2020】使用知识库嵌入改进知识图上的多跳问答

【翻译-ACL2020】使用知识库嵌入改进知识图上的多跳问答

专知会员服务

70+阅读 · 2020年7月3日

【KDD2020】图神经网络生成式预训练，GPT-GNN: Generative Pre-Training of Graph Neural Networks

【KDD2020】图神经网络生成式预训练，GPT-GNN: Generative Pre-Training of Graph Neural Networks

专知会员服务

99+阅读 · 2020年7月3日

【ACL2020】不要停止预训练:根据领域和任务自适应调整语言模型，Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

【ACL2020】不要停止预训练:根据领域和任务自适应调整语言模型，Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

专知会员服务

46+阅读 · 2020年4月25日

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

专知会员服务

96+阅读 · 2020年3月25日

【芝加哥大学】GRAPH-BERT: Only Attention is Needed for Learning Graph Representations

【芝加哥大学】GRAPH-BERT: Only Attention is Needed for Learning Graph Representations

专知会员服务

85+阅读 · 2020年1月15日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

小规模训练指南：打造世界级大语言模型的关键方法

无人机编队飞行：复杂环境中作战的策略、挑战与应用

大模型APP，AI时代第一个爆款

从数据中心视角出发的高效大语言模型训练综述

相关资讯

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

【推荐】用Tensorflow理解LSTM

【推荐】用Tensorflow理解LSTM

机器学习研究会

36+阅读 · 2017年9月11日

相关论文

Meta-Learning Online Adaptation of Language Models

Arxiv

0+阅读 · 2023年5月24日

Large Language Models are Better Reasoners with Self-Verification

Arxiv

0+阅读 · 2023年5月23日

To Copy Rather Than Memorize: A Vertical Learning Paradigm for Knowledge Graph Completion

Arxiv

0+阅读 · 2023年5月23日

Inspecting and Editing Knowledge Representations in Language Models

Arxiv

0+阅读 · 2023年5月22日

Text-to-SQL Error Correction with Language Models of Code

Arxiv

0+阅读 · 2023年5月22日

The CLIP Model is Secretly an Image-to-Prompt Converter

Arxiv

0+阅读 · 2023年5月22日

Joint Foundation Model Caching and Inference of Generative AI Services for Edge Intelligence

Arxiv

0+阅读 · 2023年5月20日

Application of Knowledge Distillation to Multi-task Speech Representation Learning

Arxiv

0+阅读 · 2023年5月19日

A One-Class Classifier for the Detection of GAN Manipulated Multi-Spectral Satellite Images

Arxiv

0+阅读 · 2023年5月19日

Which Knowledge Graph Is Best for Me?

Arxiv

11+阅读 · 2018年9月28日

相关基金

孤独症的iPSC模型研究

国家自然科学基金

1+阅读 · 2015年12月31日

儿童孤独症的神经环路机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

重组腺病毒介导Raf基因持续激活Erk1/2/Merk信号通路修复脊髓损伤的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

大白菜KIN基因的表达及其pre-mRNA加工机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

CPU/GPGPU紧耦合异构多核系统共享Last Level Cache优化研究

国家自然科学基金

0+阅读 · 2012年12月31日

金属晶粒长大动力学的多尺度模拟

国家自然科学基金

0+阅读 · 2012年12月31日

SPAC系统中农作物水循环知识融合模型研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于Linked Open Data的Web服务语义互操作关键技术

国家自然科学基金

0+阅读 · 2012年12月31日

慢病毒介导ChABC基因修饰神经干细胞治疗脊髓损伤的研究

国家自然科学基金

0+阅读 · 2011年12月31日

基因修饰的内皮祖细胞靶向治疗HER-2阳性肿瘤的实验研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员