专家的松散混集体是可通用域域学习者 (Sparse Mixture-of-Experts are Domain Generalizable Learners)

Human visual perception can easily generalize to out-of-distributed visual data, which is far beyond the capability of modern machine learning models. Domain generalization (DG) aims to close this gap, with existing DG methods mainly focusing on the loss function design. In this paper, we propose to explore an orthogonal direction, i.e., the design of the backbone architecture. It is motivated by an empirical finding that transformer-based models trained with empirical risk minimization (ERM) outperform CNN-based models employing state-of-the-art (SOTA) DG algorithms on multiple DG datasets. We develop a formal framework to characterize a network's robustness to distribution shifts by studying its architecture's alignment to the correlations in the dataset. This analysis guides us to propose a novel DG model built upon vision transformers, namely Generalizable Mixture-of-Experts (GMoE). Extensive experiments on DomainBed demonstrate that GMoE trained with ERM outperforms SOTA DG baselines by a large margin. Moreover, GMoE is complementary to existing DG methods and its performance is substantially improved when trained with DG algorithms.

翻译：人类视觉感知可以很容易地推广到分布式的视觉数据,这远远超出了现代机器学习模型的能力。域通用(DG)的目的是缩小这一差距,现有的DG方法主要侧重于损失函数设计。在本文中,我们提议探索一个正向方向,即主干结构的设计。它的动机是经验性发现,受过实验风险最小化(EMM)经验性培训的变压器模型优于以CNN为基础的模型,在多个DG数据集中采用最先进的(SOTA) DG算法。我们开发了一个正式框架,通过研究一个网络的结构与数据集的关联性,确定一个网络对分布变化的稳健性。本分析指导我们提出一个建立在视觉变异器上的新DG模型,即通用Mixture-Explects(GMOE)。关于DMeamB的大规模实验表明,经过机构风险管理培训的GDG比SOTDG基准大幅度。此外,GOE是对现有GDG方法进行重大改进时,GOE是对现有数字分析的一种补充。

相关内容

MoDELS

关注 44

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

在线变分推断，76页ppt，A Regret Bound for Online Variational Inference

专知会员服务

21+阅读 · 2019年12月2日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日