与LIMOE的多模式差异学习:专家语言图像混合 (Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts) - 专知论文

会员服务 ·

0

Learning · 混合专家模型 · 多峰值 · contrastive · Performer ·

2022 年 6 月 6 日

Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts

翻译：与LIMOE的多模式差异学习:专家语言图像混合

Basil Mustafa,Carlos Riquelme,Joan Puigcerver,Rodolphe Jenatton,Neil Houlsby

Large sparsely-activated models have obtained excellent performance in multiple domains. However, such models are typically trained on a single modality at a time. We present the Language-Image MoE, LIMoE, a sparse mixture of experts model capable of multimodal learning. LIMoE accepts both images and text simultaneously, while being trained using a contrastive loss. MoEs are a natural fit for a multimodal backbone, since expert layers can learn an appropriate partitioning of modalities. However, new challenges arise; in particular, training stability and balanced expert utilization, for which we propose an entropy-based regularization scheme. Across multiple scales, we demonstrate remarkable performance improvement over dense models of equivalent computational cost. LIMoE-L/16 trained comparably to CLIP-L/14 achieves 78.6% zero-shot ImageNet accuracy (vs. 76.2%), and when further scaled to H/14 (with additional data) it achieves 84.1%, comparable to state-of-the-art methods which use larger custom per-modality backbones and pre-training schemes. We analyse the quantitative and qualitative behavior of LIMoE, and demonstrate phenomena such as differing treatment of the modalities and the organic emergence of modality-specific experts.

翻译：然而,这些模型通常一次就单一模式进行培训。我们展示了语言图像MoE、LIMOE,这是能够多式学习的一种分散的专家模式。LIMOE同时接受图像和文本,同时接受图像和文本,同时接受对比性损失的培训。教育部自然适合多式联运主干,因为专家层可以学习适当分配模式。然而,新的挑战出现,特别是培训稳定性和平衡的专家利用,为此我们提议了一个基于酶的正规化计划。在多个尺度上,我们展示了比类似计算成本的密集模型显著的业绩改进。LIMO-L/16所培训的与CLIP-L/14相匹配的78.6%零光图像网络准确度(v. 76.2%),在进一步扩展到H/14(有额外数据)时,它达到84.1%,可与使用较大型的习惯的每个模式主干线和预培训计划相比。我们分析了LIMOE的定量和定性行为,并展示了不同模式的有机模式的出现。

1

相关内容

Learning

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

TRIB3基因表达对糖尿病大血管致纤维病变的作用及中药桃仁干预机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

中国田鼠亚科 Microtini族(Rodentia: Cricetidae: Arvicolinae)的分类与系统发育研究

国家自然科学基金

0+阅读 · 2014年12月31日

面向智能电网负荷预测的电力大数据关键技术

国家自然科学基金

8+阅读 · 2014年12月31日

PPAR β/δ基因在结直肠癌血管生成调控中的作用及分子机理

国家自然科学基金

2+阅读 · 2014年12月31日

非磁性元素掺杂稀磁半导体铁磁性机理研究的新方法

国家自然科学基金

0+阅读 · 2012年12月31日

高时效性商品在线多属性逆向拍卖定价决策与商业模式选择

国家自然科学基金

0+阅读 · 2012年12月31日

Angiopep修饰聚合物胶束的脑靶向机制及其抗HIV病毒脑内感染的研究

国家自然科学基金

0+阅读 · 2009年12月31日

超声造影微血管显像与乳腺癌血管生成和血管内皮生长因子（VEGF）表达的相关性研究

国家自然科学基金

0+阅读 · 2009年12月31日

TR3相互作用新蛋白机理研究

国家自然科学基金

1+阅读 · 2008年12月31日

TRPC6在VEGF调节新生血管形成中的作用及机制

国家自然科学基金

0+阅读 · 2008年12月31日

Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks

Arxiv

0+阅读 · 2022年7月22日

Efficient Modeling of Future Context for Image Captioning

Arxiv

0+阅读 · 2022年7月22日

Optimizing Image Compression via Joint Learning with Denoising

Arxiv

0+阅读 · 2022年7月22日

DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Arxiv

0+阅读 · 2022年7月21日

Leveraging Natural Supervision for Language Representation Learning and Generation

Arxiv

0+阅读 · 2022年7月21日

Contrastive Learning with Complex Heterogeneity

Contrastive Learning with Complex Heterogeneity

Arxiv

0+阅读 · 2022年7月21日

A Wavelet Transform and self-supervised learning-based framework for bearing fault diagnosis with limited labeled data

Arxiv

0+阅读 · 2022年7月21日

Sobolev Training for Implicit Neural Representations with Approximated Image Derivatives

Arxiv

0+阅读 · 2022年7月21日

A Simple Framework for Contrastive Learning of Visual Representations

Arxiv

21+阅读 · 2020年2月13日

Class-Balanced Loss Based on Effective Number of Samples

Arxiv

12+阅读 · 2019年1月16日

VIP会员

文章信息

相关主题

混合专家模型

相关VIP内容

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

扩散语言模型综述

《美陆军徒步机动作战条令手册》最新168页

【博士论文】理解神经网络的训练动态：从局部优化轨迹与特征学习视角

军事后勤数字化未来展望

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

相关论文

Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks

Arxiv

0+阅读 · 2022年7月22日

Efficient Modeling of Future Context for Image Captioning

Arxiv

0+阅读 · 2022年7月22日

Optimizing Image Compression via Joint Learning with Denoising

Arxiv

0+阅读 · 2022年7月22日

DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Arxiv

0+阅读 · 2022年7月21日

Leveraging Natural Supervision for Language Representation Learning and Generation

Arxiv

0+阅读 · 2022年7月21日

Contrastive Learning with Complex Heterogeneity

Contrastive Learning with Complex Heterogeneity

Arxiv

0+阅读 · 2022年7月21日

A Wavelet Transform and self-supervised learning-based framework for bearing fault diagnosis with limited labeled data

Arxiv

0+阅读 · 2022年7月21日

Sobolev Training for Implicit Neural Representations with Approximated Image Derivatives

Arxiv

0+阅读 · 2022年7月21日

A Simple Framework for Contrastive Learning of Visual Representations

Arxiv

21+阅读 · 2020年2月13日

Class-Balanced Loss Based on Effective Number of Samples

Arxiv

12+阅读 · 2019年1月16日

相关基金

TRIB3基因表达对糖尿病大血管致纤维病变的作用及中药桃仁干预机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

中国田鼠亚科 Microtini族(Rodentia: Cricetidae: Arvicolinae)的分类与系统发育研究

国家自然科学基金

0+阅读 · 2014年12月31日

面向智能电网负荷预测的电力大数据关键技术

国家自然科学基金

8+阅读 · 2014年12月31日

PPAR β/δ基因在结直肠癌血管生成调控中的作用及分子机理

国家自然科学基金

2+阅读 · 2014年12月31日

非磁性元素掺杂稀磁半导体铁磁性机理研究的新方法

国家自然科学基金

0+阅读 · 2012年12月31日

高时效性商品在线多属性逆向拍卖定价决策与商业模式选择

国家自然科学基金

0+阅读 · 2012年12月31日

Angiopep修饰聚合物胶束的脑靶向机制及其抗HIV病毒脑内感染的研究

国家自然科学基金

0+阅读 · 2009年12月31日

超声造影微血管显像与乳腺癌血管生成和血管内皮生长因子（VEGF）表达的相关性研究

国家自然科学基金

0+阅读 · 2009年12月31日

TR3相互作用新蛋白机理研究

国家自然科学基金

1+阅读 · 2008年12月31日

TRPC6在VEGF调节新生血管形成中的作用及机制

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员