一个学生知道所有专家都知道:从粗俗到浓烈 (One Student Knows All Experts Know: From Sparse to Dense) - 专知论文

会员服务 ·

0

知识 (knowledge) · 稀疏 · MoDELS · ImageNet (数据集) · 蒸馏 ·

2022 年 10 月 25 日

One Student Knows All Experts Know: From Sparse to Dense

翻译：一个学生知道所有专家都知道:从粗俗到浓烈

Fuzhao Xue,Xiaoxin He,Xiaozhe Ren,Yuxuan Lou,Yang You

Human education system trains one student by multiple experts. Mixture-of-experts (MoE) is a powerful sparse architecture including multiple experts. However, sparse MoE model is easy to overfit, hard to deploy, and not hardware-friendly for practitioners. In this work, inspired by the human education model, we propose a novel task, knowledge integration, to obtain a dense student model (OneS) as knowledgeable as one sparse MoE. We investigate this task by proposing a general training framework including knowledge gathering and knowledge distillation. Specifically, to gather key knowledge from different pre-trained experts, we first investigate four different possible knowledge gathering methods, \ie summation, averaging, Top-K Knowledge Gathering (Top-KG), and Singular Value Decomposition Knowledge Gathering (SVD-KG) proposed in this paper. We then refine the dense student model by knowledge distillation to offset the noise from gathering. On ImageNet, our OneS preserves $61.7\%$ benefits from MoE and achieves $78.4\%$ top-1 accuracy ImageNet with only $15$M parameters. On four natural language processing datasets, OneS obtains $88.2\%$ MoE benefits and outperforms the best baseline by $51.7\%$ using the same architecture and training data. In addition, compared with the MoE counterpart, OneS can achieve $3.7 \times$ inference speedup due to less computation and hardware-friendly architecture.

翻译：由多个专家培训一名学生。 Mixture of experts (MoE) 是一个强大的稀有架构,包括多位专家。然而,稀有的教育部模式很容易过度、难以部署,对实践者来说也不容易。在这项工作中,在人文教育模式的启发下,我们提出一个新的任务,即知识整合,以获得像一个稀疏的教育部那样知识丰富的学生模式。我们通过提出一个包括知识收集和知识蒸馏在内的一般培训框架来调查这项任务。具体地说,为了从不同预先培训的专家那里收集关键知识,我们首先调查四种不同的可能的知识收集方法,即: \ iet-K-KG, 平均、最高-K知识收集(Top-KG) 和 Singultanal valent Connection Connessing (SVVD-KG) 。我们然后通过知识蒸馏来改进密集的学生模式,以抵消收集的噪音。在图像网上,我们的一个S 保存$ $ $ $ $ 。我们保存了教育部的61\ $ $ $ $ $ 并实现了78.4\ $ $ $ $ 顶级图像网络, $ $ $ $ $ dirfillimmet Net et et net net net Net, et with on on on on on on on on on on cre gre gre gre gre gre gre gre gre grefrifrifrifulation, ex gregrefulation, 参数只有$$$$$$$$$ 1 参数。

0

相关内容

知识 (knowledge)

知识 (knowledge)

通过学习、实践或探索所获得的认识、判断或技能。

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

专知会员服务

30+阅读 · 2022年3月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

离子注入合成In纳米颗粒在Al薄膜中超导性质的研究

国家自然科学基金

0+阅读 · 2015年12月31日

超高频射频识别读写器芯片的多噪声建模与优化方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于灵活频谱的新型超高速光网络架构和关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

浅海低频混响中的杂波特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于混合光传输模型和复合正则化的生物发光断层成像重建方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

非局域性蒸馏

国家自然科学基金

0+阅读 · 2012年12月31日

纳米银的胚胎发育毒性机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

倏逝波放大技术实现半导体激光器超分辨率聚焦成像的研究

国家自然科学基金

0+阅读 · 2009年12月31日

联合188Re和肿瘤血管内皮特异性靶向蛋白GX/GEBP-TNF用于胃癌血管放射受体治疗

国家自然科学基金

0+阅读 · 2008年12月31日

超高性能水泥基复合材料抗多次冲击设计与动态损伤规律

国家自然科学基金

0+阅读 · 2008年12月31日

Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Arxiv

0+阅读 · 2022年12月9日

Leveraging Unlabeled Data to Track Memorization

Arxiv

0+阅读 · 2022年12月8日

Optimal Rates of Teaching and Learning Under Uncertainty

Arxiv

0+阅读 · 2022年12月8日

Predicting the Next Action by Modeling the Abstract Goal

Arxiv

0+阅读 · 2022年12月8日

Experiences from the MediaEval Predicting Media Memorability Task

Arxiv

0+阅读 · 2022年12月7日

A Survey of Learning on Small Data

Arxiv

19+阅读 · 2022年7月29日

Temporal Graph Networks for Deep Learning on Dynamic Graphs

Arxiv

37+阅读 · 2020年10月9日

Which Knowledge Graph Is Best for Me?

Arxiv

11+阅读 · 2018年9月28日

Additive Margin Softmax for Face Verification

Arxiv

11+阅读 · 2018年1月18日

The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Arxiv

11+阅读 · 2018年1月11日

VIP会员

文章信息

相关主题

知识 (knowledge)

ImageNet (数据集)

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

专知会员服务

30+阅读 · 2022年3月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美空军条令出版物：战略打击》最新条令

《高能激光武器》22页slides

军事前沿模型

《面向小型无人机或无人飞行器的创新雷达探测与人工智能分类技术》263页

相关资讯

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

相关论文

Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Arxiv

0+阅读 · 2022年12月9日

Leveraging Unlabeled Data to Track Memorization

Arxiv

0+阅读 · 2022年12月8日

Optimal Rates of Teaching and Learning Under Uncertainty

Arxiv

0+阅读 · 2022年12月8日

Predicting the Next Action by Modeling the Abstract Goal

Arxiv

0+阅读 · 2022年12月8日

Experiences from the MediaEval Predicting Media Memorability Task

Arxiv

0+阅读 · 2022年12月7日

A Survey of Learning on Small Data

Arxiv

19+阅读 · 2022年7月29日

Temporal Graph Networks for Deep Learning on Dynamic Graphs

Arxiv

37+阅读 · 2020年10月9日

Which Knowledge Graph Is Best for Me?

Arxiv

11+阅读 · 2018年9月28日

Additive Margin Softmax for Face Verification

Arxiv

11+阅读 · 2018年1月18日

The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Arxiv

11+阅读 · 2018年1月11日

相关基金

离子注入合成In纳米颗粒在Al薄膜中超导性质的研究

国家自然科学基金

0+阅读 · 2015年12月31日

超高频射频识别读写器芯片的多噪声建模与优化方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于灵活频谱的新型超高速光网络架构和关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

浅海低频混响中的杂波特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于混合光传输模型和复合正则化的生物发光断层成像重建方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

非局域性蒸馏

国家自然科学基金

0+阅读 · 2012年12月31日

纳米银的胚胎发育毒性机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

倏逝波放大技术实现半导体激光器超分辨率聚焦成像的研究

国家自然科学基金

0+阅读 · 2009年12月31日

联合188Re和肿瘤血管内皮特异性靶向蛋白GX/GEBP-TNF用于胃癌血管放射受体治疗

国家自然科学基金

0+阅读 · 2008年12月31日

超高性能水泥基复合材料抗多次冲击设计与动态损伤规律

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员