SGD噪声的校准特性及其如何帮助选择平板小型微米:稳定性分析 (The alignment property of SGD noise and how it helps select flat minima: A stability analysis) - 专知论文

会员服务 ·

0

SGD · 极小值 · 平坦最小值 · 噪声 · Frobenius 范数 ·

2022 年 10 月 17 日

The alignment property of SGD noise and how it helps select flat minima: A stability analysis

翻译：SGD噪声的校准特性及其如何帮助选择平板小型微米:稳定性分析

Lei Wu,Mingze Wang,Weijie Su

from arxiv, Accepted at NeurIPS 2022

The phenomenon that stochastic gradient descent (SGD) favors flat minima has played a critical role in understanding the implicit regularization of SGD. In this paper, we provide an explanation of this striking phenomenon by relating the particular noise structure of SGD to its \emph{linear stability} (Wu et al., 2018). Specifically, we consider training over-parameterized models with square loss. We prove that if a global minimum $\theta^*$ is linearly stable for SGD, then it must satisfy $\|H(\theta^*)\|_F\leq O(\sqrt{B}/\eta)$, where $\|H(\theta^*)\|_F, B,\eta$ denote the Frobenius norm of Hessian at $\theta^*$, batch size, and learning rate, respectively. Otherwise, SGD will escape from that minimum \emph{exponentially} fast. Hence, for minima accessible to SGD, the sharpness -- as measured by the Frobenius norm of the Hessian -- is bounded \emph{independently} of the model size and sample size. The key to obtaining these results is exploiting the particular structure of SGD noise: The noise concentrates in sharp directions of local landscape and the magnitude is proportional to loss value. This alignment property of SGD noise provably holds for linear networks and random feature models (RFMs), and is empirically verified for nonlinear networks. Moreover, the validity and practical relevance of our theoretical findings are also justified by extensive experiments on CIFAR-10 dataset.

翻译：在理解 SGD 隐含的规范化方面, SGD 偏向于平坦的梯度下降 (SGD) 现象在理解 SGD 隐含的规范化方面发挥了关键的作用。在本文中,我们通过将 SGD 的特殊噪音结构与其 emph{线性稳定性(Wu 等人, 2018) (Wu 等人, 2018) 联系起来来解释这个惊人的现象。具体地说, 我们考虑用平方损失来训练超度参数模型。否则, SGD 将很快地从最低 emph{Explential 稳定起来。因此, 它必须满足 $H(theta})\\\\\leq(leqrt{B}/\geeta) $ ($hhh(theta{theta{line sliminalityal) Oralityality reality reality reality real) 。在Hesaltimemberal ral ral deal deal deal ladeal dal dal ex lade. Srbly Srmexal ex ex ex exal deal deal deal ex ex ex exmal exmet exmal ex exmal ex ex ex ex ex ex ex ex exlev ex ex ex ex ex ex extraluttal extral. S. S. Sral exml ex ex ex ex ex ex ex ex exm ex ex ex ex extra extra extra extra extra extra ex exl exl exl ex ex ex ex ex ex ex ex ex ex ex ex exmlbal exlal ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex ex

0

相关内容

SGD

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

MIT经典《线性代数》，584页pdf，Introduction to Linear Algebra, Fifth Edition, Gilbert Strang, 2016.

MIT经典《线性代数》，584页pdf，Introduction to Linear Algebra, Fifth Edition, Gilbert Strang, 2016.

专知会员服务

428+阅读 · 2021年1月11日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

金柴抗病毒胶囊对流感病毒感染激活的GIR—I信号通路的影响研究

国家自然科学基金

0+阅读 · 2015年12月31日

NS2TP在非酒精性脂肪性肝病中保护线粒体功能的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

Heisenberg 群上的 k-平面变换

国家自然科学基金

0+阅读 · 2015年12月31日

支气管上皮细胞klotho表达在慢性阻塞性肺气肿形成中作用及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

奇性空间上的几何分析

国家自然科学基金

0+阅读 · 2013年12月31日

孤儿核受体ERRalpha作为转移性去势抵抗性前列腺癌治疗靶标的探索性研究

国家自然科学基金

0+阅读 · 2013年12月31日

A2AlTaO7基稀土钽铝酸盐设计及热物理性能调控

国家自然科学基金

0+阅读 · 2013年12月31日

低氧对大鼠EIMD肌纤维膜损伤的影响机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

关于AI-半环簇与 Conway半环簇的研究

国家自然科学基金

1+阅读 · 2012年12月31日

组合导航系统中基于混沌、小波和神经网络的信息融合方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

Sublinear-Time Computation in the Presence of Online Erasures

Arxiv

0+阅读 · 2022年11月22日

Asymptotic Properties of the Synthetic Control Method

Arxiv

0+阅读 · 2022年11月22日

Robust High-dimensional Tuning Free Multiple Testing

Arxiv

0+阅读 · 2022年11月22日

A high-order deferred correction method for the solution of free boundary problems using penalty iteration, with an application to American option pricing

Arxiv

0+阅读 · 2022年11月21日

Neural networks trained with SGD learn distributions of increasing complexity

Arxiv

0+阅读 · 2022年11月21日

Improving Sample Quality of Diffusion Models Using Self-Attention Guidance

Arxiv

0+阅读 · 2022年11月21日

Convexifying Transformers: Improving optimization and understanding of transformer networks

Arxiv

0+阅读 · 2022年11月20日

Local False Discovery Rate Based Methods for Multiple Testing of One-Way Classified Hypotheses

Arxiv

0+阅读 · 2022年11月20日

Integrating Random Effects in Deep Neural Networks

Arxiv

0+阅读 · 2022年11月20日

Explicit Second-Order Min-Max Optimization Methods with Optimal Convergence Guarantee

Arxiv

0+阅读 · 2022年11月20日

VIP会员

文章信息

相关主题

平坦最小值

Frobenius 范数

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

MIT经典《线性代数》，584页pdf，Introduction to Linear Algebra, Fifth Edition, Gilbert Strang, 2016.

MIT经典《线性代数》，584页pdf，Introduction to Linear Algebra, Fifth Edition, Gilbert Strang, 2016.

专知会员服务

428+阅读 · 2021年1月11日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

【新书】面向企业的图学习扩展：生产级图学习与推理，485页pdf

AI智能体编程：技术、挑战与机遇综述

【国家标准】数据安全技术数据安全风险评估方法

【CMU博士论文】交互式学习的进展：替代性反馈机制与自适应因果推理

相关资讯

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

相关论文

Sublinear-Time Computation in the Presence of Online Erasures

Arxiv

0+阅读 · 2022年11月22日

Asymptotic Properties of the Synthetic Control Method

Arxiv

0+阅读 · 2022年11月22日

Robust High-dimensional Tuning Free Multiple Testing

Arxiv

0+阅读 · 2022年11月22日

A high-order deferred correction method for the solution of free boundary problems using penalty iteration, with an application to American option pricing

Arxiv

0+阅读 · 2022年11月21日

Neural networks trained with SGD learn distributions of increasing complexity

Arxiv

0+阅读 · 2022年11月21日

Improving Sample Quality of Diffusion Models Using Self-Attention Guidance

Arxiv

0+阅读 · 2022年11月21日

Convexifying Transformers: Improving optimization and understanding of transformer networks

Arxiv

0+阅读 · 2022年11月20日

Local False Discovery Rate Based Methods for Multiple Testing of One-Way Classified Hypotheses

Arxiv

0+阅读 · 2022年11月20日

Integrating Random Effects in Deep Neural Networks

Arxiv

0+阅读 · 2022年11月20日

Explicit Second-Order Min-Max Optimization Methods with Optimal Convergence Guarantee

Arxiv

0+阅读 · 2022年11月20日

相关基金

金柴抗病毒胶囊对流感病毒感染激活的GIR—I信号通路的影响研究

国家自然科学基金

0+阅读 · 2015年12月31日

NS2TP在非酒精性脂肪性肝病中保护线粒体功能的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

Heisenberg 群上的 k-平面变换

国家自然科学基金

0+阅读 · 2015年12月31日

支气管上皮细胞klotho表达在慢性阻塞性肺气肿形成中作用及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

奇性空间上的几何分析

国家自然科学基金

0+阅读 · 2013年12月31日

孤儿核受体ERRalpha作为转移性去势抵抗性前列腺癌治疗靶标的探索性研究

国家自然科学基金

0+阅读 · 2013年12月31日

A2AlTaO7基稀土钽铝酸盐设计及热物理性能调控

国家自然科学基金

0+阅读 · 2013年12月31日

低氧对大鼠EIMD肌纤维膜损伤的影响机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

关于AI-半环簇与 Conway半环簇的研究

国家自然科学基金

1+阅读 · 2012年12月31日

组合导航系统中基于混沌、小波和神经网络的信息融合方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员