关于深学习模型及其内部代表的不对称问题 (On the Symmetries of Deep Learning Models and their Internal Representations)

from arxiv, CG and DB contributed equally. V2: clarified relationship between metric $\mu_{\mathrm{CKA}}$ and existing instances of CKA. V3: expanded experiment suite, alternative stitching layer capacity comparison, calculation of GeLU intertwiner group. V4: minor typo corrections

Symmetry is a fundamental tool in the exploration of a broad range of complex systems. In machine learning symmetry has been explored in both models and data. In this paper we seek to connect the symmetries arising from the architecture of a family of models with the symmetries of that family's internal representation of data. We do this by calculating a set of fundamental symmetry groups, which we call the intertwiner groups of the model. We connect intertwiner groups to a model's internal representations of data through a range of experiments that probe similarities between hidden states across models with the same architecture. Our work suggests that the symmetries of a network are propagated into the symmetries in that network's representation of data, providing us with a better understanding of how architecture affects the learning and prediction process. Finally, we speculate that for ReLU networks, the intertwiner groups may provide a justification for the common practice of concentrating model interpretability exploration on the activation basis in hidden layers rather than arbitrary linear combinations thereof.

翻译：对称是探索广泛复杂系统的基本工具。在机器学习对称中, 模型和数据都对称进行了探索。在本文中, 我们试图将模型大家庭结构产生的对称与该家庭内部数据表述的对称联系起来。我们通过计算一套基本对称组来做到这一点, 我们称之为模型的相互交错组。我们通过一系列实验将相互交错的组与模型的内部数据表述联系起来, 通过这些实验可以探测不同模型和同一结构之间的相似之处。我们的工作表明, 网络的对称在网络数据表述中的对称中传播, 使我们更好地了解结构如何影响学习和预测过程。最后, 我们推测, 对于RELU 网络, 相互交错组可能为在隐藏层而不是任意线性组合中将模型解释性探索集中在启动基础上的常见做法提供理由。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日