使用存储专家的 Tam 分散激活变形器 (Taming Sparsely Activated Transformer with Stochastic Experts) - 专知论文

会员服务 ·

0

稀疏激活 · MoDELS · 变换 · Better · BLEU ·

2022 年 2 月 3 日

Taming Sparsely Activated Transformer with Stochastic Experts

翻译：使用存储专家的 Tam 分散激活变形器

Simiao Zuo,Xiaodong Liu,Jian Jiao,Young Jin Kim,Hany Hassan,Ruofei Zhang,Tuo Zhao,Jianfeng Gao

from arxiv, ICLR 2022

Sparsely activated models (SAMs), such as Mixture-of-Experts (MoE), can easily scale to have outrageously large amounts of parameters without significant increase in computational cost. However, SAMs are reported to be parameter inefficient such that larger models do not always lead to better performance. While most on-going research focuses on improving SAMs models by exploring methods of routing inputs to experts, our analysis reveals that such research might not lead to the solution we expect, i.e., the commonly-used routing methods based on gating mechanisms do not work better than randomly routing inputs to experts. In this paper, we propose a new expert-based model, THOR (Transformer witH StOchastic ExpeRts). Unlike classic expert-based models, such as the Switch Transformer, experts in THOR are randomly activated for each input during training and inference. THOR models are trained using a consistency regularized loss, where experts learn not only from training data but also from other experts as teachers, such that all the experts make consistent predictions. We validate the effectiveness of THOR on machine translation tasks. Results show that THOR models are more parameter efficient in that they significantly outperform the Transformer and MoE models across various settings. For example, in multilingual translation, THOR outperforms the Switch Transformer by 2 BLEU scores, and obtains the same BLEU score as that of a state-of-the-art MoE model that is 18 times larger. Our code is publicly available at: https://github.com/microsoft/Stochastic-Mixture-of-Experts.

翻译：尽管大多数正在进行的研究侧重于改进SAM模型,例如Mixture Experts(MoE),可以很容易地扩大规模,从而产生令人发指的庞大参数,而不会大幅提高计算成本。然而,据报告,SAMs的参数效率低下,因此更大的模型并不总是导致更好的业绩。虽然大多数正在进行的研究侧重于通过探索向专家输入路由的方法来改进SAMs模型,但我们的分析表明,这种研究可能不会导致我们所期望的解决方案,即基于星系机制的常用变速转换方法,不会比随机地向专家输送大量的投入更好。在本文件中,我们提出了一个新的基于专家的模型,THOR(THOR) (TOR) (THOR) (THER) (THOH StOcastic ExpeRtts) (THOR) (THOR) (THOR) (THOR) (THOR) (THOR) (THOR) (ML) (MER) (MER) (MER) (TRA) (O-D) (MERL) (MERL) (OL) (OL) (OL) (OUL) (OL) (OL) (OL) (OL) (OL) (O) (OL) (O) (O) (O) (OL) (OL) (O) (O) (O) (O) (OL) (O) (O) (O) (O) (O) (O) (O) (OD) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (OD) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (OD) (OD) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (O) (

0

相关内容

稀疏激活

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

321+阅读 · 2020年11月26日

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

专知会员服务

84+阅读 · 2020年2月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

专知会员服务

24+阅读 · 2019年11月4日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新九篇自动问答相关论文—可解释推理网络、上下文知识图谱嵌入、注意力RNN、Multi-Cast注意力网络

【论文推荐】最新九篇自动问答相关论文—可解释推理网络、上下文知识图谱嵌入、注意力RNN、Multi-Cast注意力网络

专知

15+阅读 · 2018年6月29日

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

专知

20+阅读 · 2018年4月7日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

Dock3/Paks对癫痫突触可塑性的调控及异常神经网络形成机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

基于混合触发机制的CPS系统建模与分析

国家自然科学基金

0+阅读 · 2013年12月31日

基于MEMS传感器的室内个人定位技术研究

国家自然科学基金

1+阅读 · 2013年12月31日

适应婴儿肠道的双歧杆菌菌株高通量筛选体系建立

国家自然科学基金

0+阅读 · 2013年12月31日

战略性新兴产业创新生态系统的协同创新机制及其稳定性评价研究

国家自然科学基金

1+阅读 · 2013年12月31日

硫化氢在肝癌细胞乏氧辐射耐受中的作用机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

以整合酶-LEDGF/p75相互作用为靶点的抗HIV药物高通量筛选

国家自然科学基金

0+阅读 · 2012年12月31日

鸭疫里默氏杆菌整合子对其捕获和表达耐药基因盒效率的调控作用

国家自然科学基金

0+阅读 · 2012年12月31日

氢气选择性抗氧化并调节Nrf2信号转导通路抑制ROS诱导的噪声性听觉损伤的实验研究

国家自然科学基金

0+阅读 · 2011年12月31日

具有诊断与治疗多功能的超声纳米乳滴的探索性研究

国家自然科学基金

0+阅读 · 2011年12月31日

On the Representation Collapse of Sparse Mixture of Experts

Arxiv

0+阅读 · 2022年4月20日

Expert-Calibrated Learning for Online Optimization with Switching Costs

Arxiv

0+阅读 · 2022年4月18日

StableMoE: Stable Routing Strategy for Mixture of Experts

Arxiv

0+阅读 · 2022年4月18日

Distributed MST Computation in the Sleeping Model: Awake-Optimal Algorithms and Lower Bounds

Distributed MST Computation in the Sleeping Model: Awake-Optimal Algorithms and Lower Bounds

Arxiv

0+阅读 · 2022年4月18日

Inference for Cluster Randomized Experiments with Non-ignorable Cluster Sizes

Inference for Cluster Randomized Experiments with Non-ignorable Cluster Sizes

Arxiv

0+阅读 · 2022年4月18日

SkillNet: A Sparsely Activated Model for General-Purpose Natural Language Understanding

Arxiv

0+阅读 · 2022年4月18日

Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Arxiv

0+阅读 · 2022年4月16日

MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

Arxiv

0+阅读 · 2022年4月15日

Ensemble diverse hypotheses and knowledge distillation for unsupervised cross-subject adaptation

Arxiv

0+阅读 · 2022年4月15日

Analysis of Workflow Schedulers in Simulated Distributed Environments

Arxiv

0+阅读 · 2022年4月14日

VIP会员

文章信息

相关主题

相关VIP内容

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

321+阅读 · 2020年11月26日

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

专知会员服务

84+阅读 · 2020年2月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

专知会员服务

24+阅读 · 2019年11月4日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美空军条令出版物：战略打击》最新条令

《高能激光武器》22页slides

军事前沿模型

《面向小型无人机或无人飞行器的创新雷达探测与人工智能分类技术》263页

相关资讯

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新九篇自动问答相关论文—可解释推理网络、上下文知识图谱嵌入、注意力RNN、Multi-Cast注意力网络

【论文推荐】最新九篇自动问答相关论文—可解释推理网络、上下文知识图谱嵌入、注意力RNN、Multi-Cast注意力网络

专知

15+阅读 · 2018年6月29日

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

专知

20+阅读 · 2018年4月7日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

相关论文

On the Representation Collapse of Sparse Mixture of Experts

Arxiv

0+阅读 · 2022年4月20日

Expert-Calibrated Learning for Online Optimization with Switching Costs

Arxiv

0+阅读 · 2022年4月18日

StableMoE: Stable Routing Strategy for Mixture of Experts

Arxiv

0+阅读 · 2022年4月18日

Distributed MST Computation in the Sleeping Model: Awake-Optimal Algorithms and Lower Bounds

Distributed MST Computation in the Sleeping Model: Awake-Optimal Algorithms and Lower Bounds

Arxiv

0+阅读 · 2022年4月18日

Inference for Cluster Randomized Experiments with Non-ignorable Cluster Sizes

Inference for Cluster Randomized Experiments with Non-ignorable Cluster Sizes

Arxiv

0+阅读 · 2022年4月18日

SkillNet: A Sparsely Activated Model for General-Purpose Natural Language Understanding

Arxiv

0+阅读 · 2022年4月18日

Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Arxiv

0+阅读 · 2022年4月16日

MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

Arxiv

0+阅读 · 2022年4月15日

Ensemble diverse hypotheses and knowledge distillation for unsupervised cross-subject adaptation

Arxiv

0+阅读 · 2022年4月15日

Analysis of Workflow Schedulers in Simulated Distributed Environments

Arxiv

0+阅读 · 2022年4月14日

相关基金

Dock3/Paks对癫痫突触可塑性的调控及异常神经网络形成机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

基于混合触发机制的CPS系统建模与分析

国家自然科学基金

0+阅读 · 2013年12月31日

基于MEMS传感器的室内个人定位技术研究

国家自然科学基金

1+阅读 · 2013年12月31日

适应婴儿肠道的双歧杆菌菌株高通量筛选体系建立

国家自然科学基金

0+阅读 · 2013年12月31日

战略性新兴产业创新生态系统的协同创新机制及其稳定性评价研究

国家自然科学基金

1+阅读 · 2013年12月31日

硫化氢在肝癌细胞乏氧辐射耐受中的作用机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

以整合酶-LEDGF/p75相互作用为靶点的抗HIV药物高通量筛选

国家自然科学基金

0+阅读 · 2012年12月31日

鸭疫里默氏杆菌整合子对其捕获和表达耐药基因盒效率的调控作用

国家自然科学基金

0+阅读 · 2012年12月31日

氢气选择性抗氧化并调节Nrf2信号转导通路抑制ROS诱导的噪声性听觉损伤的实验研究

国家自然科学基金

0+阅读 · 2011年12月31日

具有诊断与治疗多功能的超声纳米乳滴的探索性研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员