On progressive sharpening, flat minima and generalisation - 专知论文

会员服务 ·

0

雅克比 · 平坦最小值 · 极小值 · 损失 · Networking ·

2023 年 5 月 24 日

On progressive sharpening, flat minima and generalisation

翻译：暂无翻译

Lachlan Ewen MacDonald,Jack Valmadre,Simon Lucey

We present a new approach to understanding the relationship between loss curvature and generalisation in deep learning. Specifically, we use existing empirical analyses of the spectrum of deep network loss Hessians to ground an ansatz tying together the loss Hessian and the input-output Jacobian of a deep neural network. We then prove a series of theoretical results which quantify the degree to which the input-output Jacobian of a model approximates its Lipschitz norm over a data distribution, and deduce a novel generalisation bound in terms of the empirical Jacobian. We use our ansatz, together with our theoretical results, to give a new account of the recently observed progressive sharpening phenomenon, as well as the generalisation properties of flat minima. Experimental evidence is provided to validate our claims.

翻译：暂无翻译

0

相关内容

雅克比

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】用Tensorflow理解LSTM

【推荐】用Tensorflow理解LSTM

机器学习研究会

36+阅读 · 2017年9月11日

胆盐（GCDA）诱导肝癌细胞生存与耐药的信号通路研究

国家自然科学基金

0+阅读 · 2013年12月31日

图像恢复的非局部稀疏建模理论及算法研究

国家自然科学基金

0+阅读 · 2012年12月31日

向量优化问题的近似解的最优性条件

国家自然科学基金

0+阅读 · 2012年12月31日

Stat3抑制myocardin诱导心肌肥厚的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

ERK信号转导通路在SLE表观遗传学基因表达调控机制中的作用探讨

国家自然科学基金

0+阅读 · 2012年12月31日

一类四阶MEMS方程的解集结构与解的渐近性态

国家自然科学基金

0+阅读 · 2011年12月31日

脂肪组织PTP1B表达对局部RAS的调控作用及机制研究

国家自然科学基金

0+阅读 · 2010年12月31日

我国小麦纹枯病菌Rhizoctonia cerealis的分子生态学研究

国家自然科学基金

0+阅读 · 2009年12月31日

表征团簇能量地形图拓扑结构的新方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

生物催化氧化还原途径中的遗传多样性研究

国家自然科学基金

0+阅读 · 2009年12月31日

On the hierarchical Bayesian modelling of frequency response functions

Arxiv

0+阅读 · 2023年7月12日

Implicit regularisation in stochastic gradient descent: from single-objective to two-player games

Arxiv

0+阅读 · 2023年7月11日

Fitted value shrinkage

Arxiv

0+阅读 · 2023年7月11日

Exploring Model Misspecification in Statistical Finite Elements via Shallow Water Equations

Arxiv

0+阅读 · 2023年7月11日

ProgGP: From GuitarPro Tablature Neural Generation To Progressive Metal Production

Arxiv

0+阅读 · 2023年7月11日

Prediction intervals for neural network models using weighted asymmetric loss functions

Arxiv

0+阅读 · 2023年7月11日

Entanglement Distribution in the Quantum Internet: Knowing when to Stop!

Arxiv

0+阅读 · 2023年7月11日

Cobalt: Optimizing Mining Rewards in Proof-of-Work Network Games

Arxiv

0+阅读 · 2023年7月10日

The Value of Out-of-Distribution Data

Arxiv

0+阅读 · 2023年7月10日

Continuous-Time Functional Diffusion Processes

Arxiv

0+阅读 · 2023年7月7日

VIP会员

文章信息

相关主题

平坦最小值

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《代码、指挥与冲突：描绘军事人工智能的未来》报告

【斯坦福博士论文】面向地理空间数据的多模态与多尺度建模：时空生成式人工智能

美国启动“自有军事人工智能计划”：采用谷歌Gemini以推动全军人工智能应用

《创新与适应性作为军事成功的关键因素：来自俄乌战争的战略洞见》报告

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】用Tensorflow理解LSTM

【推荐】用Tensorflow理解LSTM

机器学习研究会

36+阅读 · 2017年9月11日

相关论文

On the hierarchical Bayesian modelling of frequency response functions

Arxiv

0+阅读 · 2023年7月12日

Implicit regularisation in stochastic gradient descent: from single-objective to two-player games

Arxiv

0+阅读 · 2023年7月11日

Fitted value shrinkage

Arxiv

0+阅读 · 2023年7月11日

Exploring Model Misspecification in Statistical Finite Elements via Shallow Water Equations

Arxiv

0+阅读 · 2023年7月11日

ProgGP: From GuitarPro Tablature Neural Generation To Progressive Metal Production

Arxiv

0+阅读 · 2023年7月11日

Prediction intervals for neural network models using weighted asymmetric loss functions

Arxiv

0+阅读 · 2023年7月11日

Entanglement Distribution in the Quantum Internet: Knowing when to Stop!

Arxiv

0+阅读 · 2023年7月11日

Cobalt: Optimizing Mining Rewards in Proof-of-Work Network Games

Arxiv

0+阅读 · 2023年7月10日

The Value of Out-of-Distribution Data

Arxiv

0+阅读 · 2023年7月10日

Continuous-Time Functional Diffusion Processes

Arxiv

0+阅读 · 2023年7月7日

相关基金

胆盐（GCDA）诱导肝癌细胞生存与耐药的信号通路研究

国家自然科学基金

0+阅读 · 2013年12月31日

图像恢复的非局部稀疏建模理论及算法研究

国家自然科学基金

0+阅读 · 2012年12月31日

向量优化问题的近似解的最优性条件

国家自然科学基金

0+阅读 · 2012年12月31日

Stat3抑制myocardin诱导心肌肥厚的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

ERK信号转导通路在SLE表观遗传学基因表达调控机制中的作用探讨

国家自然科学基金

0+阅读 · 2012年12月31日

一类四阶MEMS方程的解集结构与解的渐近性态

国家自然科学基金

0+阅读 · 2011年12月31日

脂肪组织PTP1B表达对局部RAS的调控作用及机制研究

国家自然科学基金

0+阅读 · 2010年12月31日

我国小麦纹枯病菌Rhizoctonia cerealis的分子生态学研究

国家自然科学基金

0+阅读 · 2009年12月31日

表征团簇能量地形图拓扑结构的新方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

生物催化氧化还原途径中的遗传多样性研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员