MLP-Mixer as a Wide and Sparse MLP - 专知论文

会员服务 ·

0

稀疏 · 层 · 混合 · Weight · 宽度 ·

2023 年 6 月 2 日

MLP-Mixer as a Wide and Sparse MLP

翻译：暂无翻译

Tomohiro Hayase,Ryo Karakida

from arxiv, 19 pages, 13 figures

Multi-layer perceptron (MLP) is a fundamental component of deep learning that has been extensively employed for various problems. However, recent empirical successes in MLP-based architectures, particularly the progress of the MLP-Mixer, have revealed that there is still hidden potential in improving MLPs to achieve better performance. In this study, we reveal that the MLP-Mixer works effectively as a wide MLP with certain sparse weights. Initially, we clarify that the mixing layer of the Mixer has an effective expression as a wider MLP whose weights are sparse and represented by the Kronecker product. This expression naturally defines a permuted-Kronecker (PK) family, which can be regarded as a general class of mixing layers and is also regarded as an approximation of Monarch matrices. Subsequently, because the PK family effectively constitutes a wide MLP with sparse weights, one can apply the hypothesis proposed by Golubeva, Neyshabur and Gur-Ari (2021) that the prediction performance improves as the width (sparsity) increases when the number of weights is fixed. We empirically verify this hypothesis by maximizing the effective width of the MLP-Mixer, which enables us to determine the appropriate size of the mixing layers quantitatively.

翻译：暂无翻译

0

相关内容

CVPR 2023开会了！谷歌等最新《视觉上理解和解释注意力》教程，附152页ppt

CVPR 2023开会了！谷歌等最新《视觉上理解和解释注意力》教程，附152页ppt

专知会员服务

85+阅读 · 2023年6月19日

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

2021机器学习研究风向是啥？MLP→CNN→Transformer→MLP！

2021机器学习研究风向是啥？MLP→CNN→Transformer→MLP！

专知会员服务

67+阅读 · 2021年5月23日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

深度自进化聚类：Deep Self-Evolution Clustering

深度自进化聚类：Deep Self-Evolution Clustering

我爱读PAMI

15+阅读 · 2019年4月13日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

【推荐】YOLO实时目标检测(6fps)

【推荐】YOLO实时目标检测(6fps)

机器学习研究会

20+阅读 · 2017年11月5日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

【推荐】图像分类必读开创性论文汇总

【推荐】图像分类必读开创性论文汇总

机器学习研究会

14+阅读 · 2017年8月15日

拓扑绝缘体/Si异质结能带调控与器件应用基础研究

国家自然科学基金

0+阅读 · 2014年12月31日

影响东亚冬季气候的海温和北极海冰配置型

国家自然科学基金

0+阅读 · 2013年12月31日

半导体衬底上FeSe薄膜的外延生长及界面超导

国家自然科学基金

0+阅读 · 2013年12月31日

4f和3d电子调控下的新型In和Te基稀土1：3型半导体化合物的磁输运和结构

国家自然科学基金

0+阅读 · 2012年12月31日

SND1蛋白与PML蛋白相互作用在APL中促进白血病细胞增殖和抑制分化作用机制的研究

国家自然科学基金

0+阅读 · 2012年12月31日

Anxa2介导IL-6诱导的SHP2/Erk和JAK2/STAT3信号通路激活并促进乳腺癌细胞上皮间质转化的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Snail1调控STOML2的表达在糖尿病肾病EMT发生中的作用及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

钙钛矿氧化物薄膜异质界面的奇异磁性和磁输运

国家自然科学基金

0+阅读 · 2011年12月31日

长链非编码RNA在急性髓系白血病t(8;21)和inv(16)型的调控作用及其机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

SARI基因在肺癌侵袭转移中的作用及分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

Exphormer: Sparse Transformers for Graphs

Arxiv

0+阅读 · 2023年7月24日

Dropout Drops Double Descent

Arxiv

0+阅读 · 2023年7月22日

Representation Power of Graph Neural Networks: Improved Expressivity via Algebraic Analysis

Arxiv

0+阅读 · 2023年7月21日

On the Universality of Linear Recurrences Followed by Nonlinear Projections

Arxiv

0+阅读 · 2023年7月21日

On the power of counting the total number of computation paths of NPTMs

Arxiv

0+阅读 · 2023年7月21日

Spectral Universality of Regularized Linear Regression with Nearly Deterministic Sensing Matrices

Arxiv

0+阅读 · 2023年7月20日

Here Comes the STRAIN: Analyzing Defensive Pass Rush in American Football with Player Tracking Data

Arxiv

0+阅读 · 2023年7月20日

High-order Tensor Pooling with Attention for Action Recognition

Arxiv

0+阅读 · 2023年7月20日

Consistent Group selection using Global-local prior in High dimensional setup

Arxiv

0+阅读 · 2023年7月20日

A test for counting sequences of integer-valued autoregressive models

Arxiv

0+阅读 · 2023年7月18日

VIP会员

文章信息

相关主题

相关VIP内容

CVPR 2023开会了！谷歌等最新《视觉上理解和解释注意力》教程，附152页ppt

CVPR 2023开会了！谷歌等最新《视觉上理解和解释注意力》教程，附152页ppt

专知会员服务

85+阅读 · 2023年6月19日

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

2021机器学习研究风向是啥？MLP→CNN→Transformer→MLP！

2021机器学习研究风向是啥？MLP→CNN→Transformer→MLP！

专知会员服务

67+阅读 · 2021年5月23日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津大学博士论文】将序列结构与几何结构融入深度神经网络

工程视角：影响战争进程的小型无人机

企业级AI应用开发：从技术选型到生产落地

AI生成代码缺陷综述

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

深度自进化聚类：Deep Self-Evolution Clustering

深度自进化聚类：Deep Self-Evolution Clustering

我爱读PAMI

15+阅读 · 2019年4月13日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

【推荐】YOLO实时目标检测(6fps)

【推荐】YOLO实时目标检测(6fps)

机器学习研究会

20+阅读 · 2017年11月5日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

【推荐】图像分类必读开创性论文汇总

【推荐】图像分类必读开创性论文汇总

机器学习研究会

14+阅读 · 2017年8月15日

相关论文

Exphormer: Sparse Transformers for Graphs

Arxiv

0+阅读 · 2023年7月24日

Dropout Drops Double Descent

Arxiv

0+阅读 · 2023年7月22日

Representation Power of Graph Neural Networks: Improved Expressivity via Algebraic Analysis

Arxiv

0+阅读 · 2023年7月21日

On the Universality of Linear Recurrences Followed by Nonlinear Projections

Arxiv

0+阅读 · 2023年7月21日

On the power of counting the total number of computation paths of NPTMs

Arxiv

0+阅读 · 2023年7月21日

Spectral Universality of Regularized Linear Regression with Nearly Deterministic Sensing Matrices

Arxiv

0+阅读 · 2023年7月20日

Here Comes the STRAIN: Analyzing Defensive Pass Rush in American Football with Player Tracking Data

Arxiv

0+阅读 · 2023年7月20日

High-order Tensor Pooling with Attention for Action Recognition

Arxiv

0+阅读 · 2023年7月20日

Consistent Group selection using Global-local prior in High dimensional setup

Arxiv

0+阅读 · 2023年7月20日

A test for counting sequences of integer-valued autoregressive models

Arxiv

0+阅读 · 2023年7月18日

相关基金

拓扑绝缘体/Si异质结能带调控与器件应用基础研究

国家自然科学基金

0+阅读 · 2014年12月31日

影响东亚冬季气候的海温和北极海冰配置型

国家自然科学基金

0+阅读 · 2013年12月31日

半导体衬底上FeSe薄膜的外延生长及界面超导

国家自然科学基金

0+阅读 · 2013年12月31日

4f和3d电子调控下的新型In和Te基稀土1：3型半导体化合物的磁输运和结构

国家自然科学基金

0+阅读 · 2012年12月31日

SND1蛋白与PML蛋白相互作用在APL中促进白血病细胞增殖和抑制分化作用机制的研究

国家自然科学基金

0+阅读 · 2012年12月31日

Anxa2介导IL-6诱导的SHP2/Erk和JAK2/STAT3信号通路激活并促进乳腺癌细胞上皮间质转化的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Snail1调控STOML2的表达在糖尿病肾病EMT发生中的作用及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

钙钛矿氧化物薄膜异质界面的奇异磁性和磁输运

国家自然科学基金

0+阅读 · 2011年12月31日

长链非编码RNA在急性髓系白血病t(8;21)和inv(16)型的调控作用及其机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

SARI基因在肺癌侵袭转移中的作用及分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员