特派专家特派专家 (Task-Specific Expert Pruning for Sparse Mixture-of-Experts) - 专知论文

会员服务 ·

0

稀疏 · MoDELS · 推断 · 可约的 · 剪枝 ·

2022 年 6 月 1 日

Task-Specific Expert Pruning for Sparse Mixture-of-Experts

翻译：特派专家特派专家

Tianyu Chen,Shaohan Huang,Yuan Xie,Binxing Jiao,Daxin Jiang,Haoyi Zhou,Jianxin Li,Furu Wei

from arxiv, under review

The sparse Mixture-of-Experts (MoE) model is powerful for large-scale pre-training and has achieved promising results due to its model capacity. However, with trillions of parameters, MoE is hard to be deployed on cloud or mobile environment. The inference of MoE requires expert parallelism, which is not hardware-friendly and communication expensive. Especially for resource-limited downstream tasks, such sparse structure has to sacrifice a lot of computing efficiency for limited performance gains. In this work, we observe most experts contribute scarcely little to the MoE fine-tuning and inference. We further propose a general method to progressively drop the non-professional experts for the target downstream task, which preserves the benefits of MoE while reducing the MoE model into one single-expert dense model. Our experiments reveal that the fine-tuned single-expert model could preserve 99.3% benefits from MoE across six different types of tasks while enjoying 2x inference speed with free communication cost.

翻译：稀有的专家混合(Mixture of Experters)模式对于大规模培训前的训练十分强大,并因其模型能力而取得了可喜的成果。然而,由于有数万亿参数,教育部很难在云层或移动环境中部署。教育部的推论要求专家平行,这不易硬件使用,通信费用昂贵。特别是对于资源有限的下游任务,这种稀有的结构必须牺牲大量计算效率,以取得有限的绩效收益。在这项工作中,我们观察到大多数专家对教育部的微调和推论贡献很少。我们进一步提出了逐步减少非专业专家从事目标下游任务的一般方法,这种方法既维护教育部模式的好处,又将教育部模式减为单一的专家密集模式。我们的实验表明,微调的单一专家模式可以保存教育部在六种不同任务中的99.3%的利益,同时享有2x的免费通信成本的推断速度。

0

相关内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

miR-506多靶点调控HR和β-catenin信号通路对浆液性卵巢癌药物敏感性的影响

国家自然科学基金

0+阅读 · 2014年12月31日

以分泌型热休克蛋白90α为靶标的抗癌化合物的设计、合成和活性评价

国家自然科学基金

0+阅读 · 2014年12月31日

PPAR β/δ基因在结直肠癌血管生成调控中的作用及分子机理

国家自然科学基金

2+阅读 · 2014年12月31日

TRPV4在Aβ诱导星形胶质细胞活化及介导神经元死亡中的作用

国家自然科学基金

0+阅读 · 2013年12月31日

功能化的氧化石墨烯诱导沸石合成及表面负载

国家自然科学基金

0+阅读 · 2013年12月31日

拟南芥光敏色素A的蛋白磷酸化调控其信号传导的分子机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

microRNAs在MLCK调控动脉粥样硬化血管重构中的作用及分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

mTOR信号通路在痛相关海马突触可塑性中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

nanog对牙髓干细胞增殖分化的影响及信号通路调控

国家自然科学基金

0+阅读 · 2011年12月31日

以热休克蛋白90为靶标的抗癌化合物的设计、合成和活性评价

国家自然科学基金

1+阅读 · 2009年12月31日

On the Usability of Transformers-based models for a French Question-Answering task

Arxiv

0+阅读 · 2022年7月19日

MoEC: Mixture of Expert Clusters

Arxiv

0+阅读 · 2022年7月19日

Robust Training of Neural Networks Using Scale Invariant Architectures

Arxiv

0+阅读 · 2022年7月18日

Improved optimization strategies for deep Multi-Task Networks

Arxiv

0+阅读 · 2022年7月18日

The Multiple Subnetwork Hypothesis: Enabling Multidomain Learning by Isolating Task-Specific Subnetworks in Feedforward Neural Networks

Arxiv

0+阅读 · 2022年7月18日

Comprehensive Graph Gradual Pruning for Sparse Training in Graph Neural Networks

Arxiv

0+阅读 · 2022年7月18日

Understanding the Generalization Performance of Spectral Clustering Algorithms

Arxiv

0+阅读 · 2022年7月17日

The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification Challenge

The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification Challenge

Arxiv

0+阅读 · 2022年7月15日

Prompt Injection: Parameterization of Fixed Inputs

Arxiv

1+阅读 · 2022年7月15日

A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP

Arxiv

12+阅读 · 2021年8月30日

VIP会员

文章信息

相关主题

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军特种作战条令》最新102页

《洛克希德SR-71“黑鸟”侦察机动力系统》21页slides

美空军作战实验室通过人工智能和指挥控制技术创新推进杀伤链

《指挥控制能力分析方法论》最新报告

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

On the Usability of Transformers-based models for a French Question-Answering task

Arxiv

0+阅读 · 2022年7月19日

MoEC: Mixture of Expert Clusters

Arxiv

0+阅读 · 2022年7月19日

Robust Training of Neural Networks Using Scale Invariant Architectures

Arxiv

0+阅读 · 2022年7月18日

Improved optimization strategies for deep Multi-Task Networks

Arxiv

0+阅读 · 2022年7月18日

The Multiple Subnetwork Hypothesis: Enabling Multidomain Learning by Isolating Task-Specific Subnetworks in Feedforward Neural Networks

Arxiv

0+阅读 · 2022年7月18日

Comprehensive Graph Gradual Pruning for Sparse Training in Graph Neural Networks

Arxiv

0+阅读 · 2022年7月18日

Understanding the Generalization Performance of Spectral Clustering Algorithms

Arxiv

0+阅读 · 2022年7月17日

The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification Challenge

The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification Challenge

Arxiv

0+阅读 · 2022年7月15日

Prompt Injection: Parameterization of Fixed Inputs

Arxiv

1+阅读 · 2022年7月15日

A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP

Arxiv

12+阅读 · 2021年8月30日

相关基金

miR-506多靶点调控HR和β-catenin信号通路对浆液性卵巢癌药物敏感性的影响

国家自然科学基金

0+阅读 · 2014年12月31日

以分泌型热休克蛋白90α为靶标的抗癌化合物的设计、合成和活性评价

国家自然科学基金

0+阅读 · 2014年12月31日

PPAR β/δ基因在结直肠癌血管生成调控中的作用及分子机理

国家自然科学基金

2+阅读 · 2014年12月31日

TRPV4在Aβ诱导星形胶质细胞活化及介导神经元死亡中的作用

国家自然科学基金

0+阅读 · 2013年12月31日

功能化的氧化石墨烯诱导沸石合成及表面负载

国家自然科学基金

0+阅读 · 2013年12月31日

拟南芥光敏色素A的蛋白磷酸化调控其信号传导的分子机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

microRNAs在MLCK调控动脉粥样硬化血管重构中的作用及分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

mTOR信号通路在痛相关海马突触可塑性中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

nanog对牙髓干细胞增殖分化的影响及信号通路调控

国家自然科学基金

0+阅读 · 2011年12月31日

以热休克蛋白90为靶标的抗癌化合物的设计、合成和活性评价

国家自然科学基金

1+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员