从块Toeplitz矩阵到图上微分方程: 朝着可扩展掩码Transformer的一般理论 (From block-Toeplitz matrices to differential equations on graphs: towards a general theory for scalable masked Transformers) - 专知论文

会员服务 ·

0

掩码 · Toeplitz矩阵 · Transformer · 谱分析 · 马尔科夫 ·

2023 年 3 月 28 日

From block-Toeplitz matrices to differential equations on graphs: towards a general theory for scalable masked Transformers

翻译：从块Toeplitz矩阵到图上微分方程: 朝着可扩展掩码Transformer的一般理论

Krzysztof Choromanski,Han Lin,Haoxian Chen,Tianyi Zhang,Arijit Sehanobish,Valerii Likhosherstov,Jack Parker-Holder,Tamas Sarlos,Adrian Weller,Thomas Weingarten

from arxiv, 20 pages, 12 figures

In this paper we provide, to the best of our knowledge, the first comprehensive approach for incorporating various masking mechanisms into Transformers architectures in a scalable way. We show that recent results on linear causal attention (Choromanski et al., 2021) and log-linear RPE-attention (Luo et al., 2021) are special cases of this general mechanism. However by casting the problem as a topological (graph-based) modulation of unmasked attention, we obtain several results unknown before, including efficient d-dimensional RPE-masking and graph-kernel masking. We leverage many mathematical techniques ranging from spectral analysis through dynamic programming and random walks to new algorithms for solving Markov processes on graphs. We provide a corresponding empirical evaluation.

翻译：在本文中，我们提供了一个全面的方法，以可扩展的方式将各种掩码机制合并到Transformer体系结构中。我们展示了最近关于线性因果注意力（Choromanski等人，2021）和对数线性RPE-注意力（Luo等人，2021）的结果是这种通用机制的特例。但是，通过将问题视为未掩码注意力的拓扑（基于图形）调制，我们获得了多个以前不为人知的结果，包括高效的d维RPE掩码和图内核掩码。我们利用许多数学技术，从谱分析到动态规划和随机游走，以及解决图上马尔科夫过程的新算法。我们提供相应的实验证明。

0

相关内容

神经网络数学基础，45页ppt

神经网络数学基础，45页ppt

专知会员服务

83+阅读 · 2023年5月7日

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

专知会员服务

18+阅读 · 2022年8月8日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【硬核书】矩阵代数基础，248页pdf

【硬核书】矩阵代数基础，248页pdf

专知会员服务

87+阅读 · 2021年12月9日

【CVPR2021】用Transformers无监督预训练进行目标检测

【CVPR2021】用Transformers无监督预训练进行目标检测

专知会员服务

58+阅读 · 2021年3月3日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

320+阅读 · 2020年11月26日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

康奈尔大学Jon Kleinberg经典书《算法设计Algorithm Design》课件PPT与电子书，864页pdf

康奈尔大学Jon Kleinberg经典书《算法设计Algorithm Design》课件PPT与电子书，864页pdf

专知会员服务

235+阅读 · 2020年1月21日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

专知

0+阅读 · 2022年8月8日

使用BERT做文本摘要

使用BERT做文本摘要

专知

23+阅读 · 2019年12月7日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

专知

14+阅读 · 2018年3月30日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】GAN架构入门综述(资源汇总)

【推荐】GAN架构入门综述(资源汇总)

机器学习研究会

10+阅读 · 2017年9月3日

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

Heisenberg群与Minkowski空间中的非线性椭圆方程

国家自然科学基金

0+阅读 · 2014年12月31日

冷原子系统中的可积模型

国家自然科学基金

0+阅读 · 2013年12月31日

系数变号的广义Abel方程的几何性质与周期解问题

国家自然科学基金

0+阅读 · 2013年12月31日

EAST装置上低杂波电流驱动对2/1新经典撕裂模的致稳研究

国家自然科学基金

0+阅读 · 2012年12月31日

偏微分方程的正则性

国家自然科学基金

3+阅读 · 2012年12月31日

有机分子半导体的非局域电声子耦合：声子色散与二阶电声子相互作用的影响

国家自然科学基金

0+阅读 · 2012年12月31日

非经典耗散方程的谱方法与动力学研究

国家自然科学基金

1+阅读 · 2011年12月31日

线性积分方程的Galerkin快速谱方法

国家自然科学基金

0+阅读 · 2009年12月31日

约化群酉表示的branching law及其应用

国家自然科学基金

0+阅读 · 2009年12月31日

Using Symbolic Computation to Analyze Zero-Hopf Bifurcations of Polynomial Differential Systems

Arxiv

0+阅读 · 2023年5月18日

Dynamic Matrix Recovery

Arxiv

0+阅读 · 2023年5月17日

Value Iteration Networks with Gated Summarization Module

Arxiv

0+阅读 · 2023年5月16日

Finding Regions of Counterfactual Explanations via Robust Optimization

Arxiv

0+阅读 · 2023年5月16日

Preferential Pliable Index Coding

Arxiv

0+阅读 · 2023年5月15日

SKI to go Faster: Accelerating Toeplitz Neural Networks via Asymmetric Kernels

Arxiv

0+阅读 · 2023年5月15日

Graph Ordering Attention Networks

Arxiv

12+阅读 · 2022年11月21日

Transformers in Time Series: A Survey

Arxiv

34+阅读 · 2022年2月15日

Unifying Graph Convolutional Neural Networks and Label Propagation

Arxiv

31+阅读 · 2020年2月17日

Hierarchical Graph Representation Learning with Differentiable Pooling

Hierarchical Graph Representation Learning with Differentiable Pooling

Arxiv

13+阅读 · 2018年6月26日

VIP会员

文章信息

相关主题

相关VIP内容

神经网络数学基础，45页ppt

神经网络数学基础，45页ppt

专知会员服务

83+阅读 · 2023年5月7日

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

专知会员服务

18+阅读 · 2022年8月8日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【硬核书】矩阵代数基础，248页pdf

【硬核书】矩阵代数基础，248页pdf

专知会员服务

87+阅读 · 2021年12月9日

【CVPR2021】用Transformers无监督预训练进行目标检测

【CVPR2021】用Transformers无监督预训练进行目标检测

专知会员服务

58+阅读 · 2021年3月3日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

320+阅读 · 2020年11月26日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

康奈尔大学Jon Kleinberg经典书《算法设计Algorithm Design》课件PPT与电子书，864页pdf

康奈尔大学Jon Kleinberg经典书《算法设计Algorithm Design》课件PPT与电子书，864页pdf

专知会员服务

235+阅读 · 2020年1月21日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

热门VIP内容

开通专知VIP会员享更多权益服务

《基于大型语言模型的软件工程自动化研究》最新264页

《基于大型语言模型的信号处理管线研究：推进军事电子情报工作流程》最新76页

中文版 | 战争算法：生成式人工智能在战场的崛起

中文版《美国陆军：战术行为性远程医疗实施观察与建议》

相关资讯

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

【ICML2022】从block-Toeplitz矩阵到图上的微分方程:迈向可扩展掩码Transformers的一般理论

专知

0+阅读 · 2022年8月8日

使用BERT做文本摘要

使用BERT做文本摘要

专知

23+阅读 · 2019年12月7日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

专知

14+阅读 · 2018年3月30日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】GAN架构入门综述(资源汇总)

【推荐】GAN架构入门综述(资源汇总)

机器学习研究会

10+阅读 · 2017年9月3日

相关论文

Using Symbolic Computation to Analyze Zero-Hopf Bifurcations of Polynomial Differential Systems

Arxiv

0+阅读 · 2023年5月18日

Dynamic Matrix Recovery

Arxiv

0+阅读 · 2023年5月17日

Value Iteration Networks with Gated Summarization Module

Arxiv

0+阅读 · 2023年5月16日

Finding Regions of Counterfactual Explanations via Robust Optimization

Arxiv

0+阅读 · 2023年5月16日

Preferential Pliable Index Coding

Arxiv

0+阅读 · 2023年5月15日

SKI to go Faster: Accelerating Toeplitz Neural Networks via Asymmetric Kernels

Arxiv

0+阅读 · 2023年5月15日

Graph Ordering Attention Networks

Arxiv

12+阅读 · 2022年11月21日

Transformers in Time Series: A Survey

Arxiv

34+阅读 · 2022年2月15日

Unifying Graph Convolutional Neural Networks and Label Propagation

Arxiv

31+阅读 · 2020年2月17日

Hierarchical Graph Representation Learning with Differentiable Pooling

Hierarchical Graph Representation Learning with Differentiable Pooling

Arxiv

13+阅读 · 2018年6月26日

相关基金

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

Heisenberg群与Minkowski空间中的非线性椭圆方程

国家自然科学基金

0+阅读 · 2014年12月31日

冷原子系统中的可积模型

国家自然科学基金

0+阅读 · 2013年12月31日

系数变号的广义Abel方程的几何性质与周期解问题

国家自然科学基金

0+阅读 · 2013年12月31日

EAST装置上低杂波电流驱动对2/1新经典撕裂模的致稳研究

国家自然科学基金

0+阅读 · 2012年12月31日

偏微分方程的正则性

国家自然科学基金

3+阅读 · 2012年12月31日

有机分子半导体的非局域电声子耦合：声子色散与二阶电声子相互作用的影响

国家自然科学基金

0+阅读 · 2012年12月31日

非经典耗散方程的谱方法与动力学研究

国家自然科学基金

1+阅读 · 2011年12月31日

线性积分方程的Galerkin快速谱方法

国家自然科学基金

0+阅读 · 2009年12月31日

约化群酉表示的branching law及其应用

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员