由GD培训的浅水超度参数化神经网络的无对数回归,早期停止 (Nonparametric Regression with Shallow Overparameterized Neural Networks Trained by GD with Early Stopping) - 专知论文

会员服务 ·

0

早停 · Neural Networks · Networking · CASE · 噪声 ·

2021 年 7 月 12 日

Nonparametric Regression with Shallow Overparameterized Neural Networks Trained by GD with Early Stopping

翻译：由GD培训的浅水超度参数化神经网络的无对数回归,早期停止

Ilja Kuzborskij,Csaba Szepesvári

from arxiv, COLT 2021

We explore the ability of overparameterized shallow neural networks to learn Lipschitz regression functions with and without label noise when trained by Gradient Descent (GD). To avoid the problem that in the presence of noisy labels, neural networks trained to nearly zero training error are inconsistent on this class, we propose an early stopping rule that allows us to show optimal rates. This provides an alternative to the result of Hu et al. (2021) who studied the performance of $\ell 2$ -regularized GD for training shallow networks in nonparametric regression which fully relied on the infinite-width network (Neural Tangent Kernel (NTK)) approximation. Here we present a simpler analysis which is based on a partitioning argument of the input space (as in the case of 1-nearest-neighbor rule) coupled with the fact that trained neural networks are smooth with respect to their inputs when trained by GD. In the noise-free case the proof does not rely on any kernelization and can be regarded as a finite-width result. In the case of label noise, by slightly modifying the proof, the noise is controlled using a technique of Yao, Rosasco, and Caponnetto (2007).

翻译：我们探讨过量的浅浅神经网络是否有能力在受Gradient Emproper(GD)培训时,学习使用和不带标签噪音的Lipschitz回归功能。为了避免在出现噪音标签的情况下,经过训练的神经网络在这个班级上几乎是零培训错误的问题不一致,我们提议了一项早期停止规则,使我们能够显示最佳比率。这为Hu等人(2021年)研究了2美元-正规化GD的性能,以培训完全依赖无限宽网(Neural Tangent Kernel(NTKKK)))近似的非临界回归的浅网络提供了一种替代方法。在这里,我们提出一项更简单的分析,其依据是对输入空间进行分割的争论(如1个近邻邻居规则),以及经过训练的神经网络在接受GD(2021年)培训时其投入方面是顺畅的。在无噪音的情况下,证据并不依赖任何内核化,而且可以被视为一种限定的边缘结果。在标签噪音的情况下,通过稍微修改证据,将噪音控制起来,并使用Casion 和Yasco (2007年) 技术。

0

相关内容

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

神经网络序列数据建模，229页ppt，Modeling Sequential Data with Neural Nets

神经网络序列数据建模，229页ppt，Modeling Sequential Data with Neural Nets

专知会员服务

67+阅读 · 2020年7月25日

【Google】平滑对抗训练，Smooth Adversarial Training

【Google】平滑对抗训练，Smooth Adversarial Training

专知会员服务

49+阅读 · 2020年7月4日

Fariz Darari简明《博弈论Game Theory》介绍，35页ppt

Fariz Darari简明《博弈论Game Theory》介绍，35页ppt

专知会员服务

112+阅读 · 2020年5月15日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【Google】神经架构搜索（Neural Architecture Search and Beyond），Barret Zoph

【Google】神经架构搜索（Neural Architecture Search and Beyond），Barret Zoph

专知会员服务

31+阅读 · 2019年11月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

已删除

inpluslab

8+阅读 · 2019年10月29日

【学习】Hierarchical Softmax

【学习】Hierarchical Softmax

机器学习研究会

4+阅读 · 2017年8月6日

Uniform Generalization Bounds for Overparameterized Neural Networks

Uniform Generalization Bounds for Overparameterized Neural Networks

Arxiv

0+阅读 · 2021年9月13日

Optimal Classification for Functional Data

Arxiv

0+阅读 · 2021年9月10日

Sharing Matters for Generalization in Deep Metric Learning

Arxiv

0+阅读 · 2021年9月9日

Sharp Lower Bounds on the Approximation Rate of Shallow Neural Networks

Arxiv

0+阅读 · 2021年9月8日

Wave-Informed Matrix Factorization with Global Optimality Guarantees

Arxiv

0+阅读 · 2021年9月8日

Scaling Properties of Deep Residual Networks

Arxiv

13+阅读 · 2021年5月25日

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Arxiv

13+阅读 · 2020年6月24日

Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Arxiv

8+阅读 · 2018年11月21日

Reducing Parameter Space for Neural Network Training

Arxiv

3+阅读 · 2018年8月17日

Classification with Fairness Constraints: A Meta-Algorithm with Provable Guarantees

Classification with Fairness Constraints: A Meta-Algorithm with Provable Guarantees

Arxiv

3+阅读 · 2018年8月2日

VIP会员

文章信息

相关主题

Neural Networks

相关VIP内容

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

神经网络序列数据建模，229页ppt，Modeling Sequential Data with Neural Nets

神经网络序列数据建模，229页ppt，Modeling Sequential Data with Neural Nets

专知会员服务

67+阅读 · 2020年7月25日

【Google】平滑对抗训练，Smooth Adversarial Training

【Google】平滑对抗训练，Smooth Adversarial Training

专知会员服务

49+阅读 · 2020年7月4日

Fariz Darari简明《博弈论Game Theory》介绍，35页ppt

Fariz Darari简明《博弈论Game Theory》介绍，35页ppt

专知会员服务

112+阅读 · 2020年5月15日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【Google】神经架构搜索（Neural Architecture Search and Beyond），Barret Zoph

【Google】神经架构搜索（Neural Architecture Search and Beyond），Barret Zoph

专知会员服务

31+阅读 · 2019年11月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

【ACMMM2025教程】打击网络虚假信息视频：特征分析、检测与防范，170页ppt

海军无人系统：海上作战的演进而非革命

Nature 子刊 | SciToolAgent:知识图谱引导的科学工具智能体

多媒体顶会ACM Multimedia 2025各大奖项揭晓！格拉斯哥大学等获最佳论文，中科院自动化所等获最佳学生论文

相关资讯

已删除

inpluslab

8+阅读 · 2019年10月29日

【学习】Hierarchical Softmax

【学习】Hierarchical Softmax

机器学习研究会

4+阅读 · 2017年8月6日

相关论文

Uniform Generalization Bounds for Overparameterized Neural Networks

Uniform Generalization Bounds for Overparameterized Neural Networks

Arxiv

0+阅读 · 2021年9月13日

Optimal Classification for Functional Data

Arxiv

0+阅读 · 2021年9月10日

Sharing Matters for Generalization in Deep Metric Learning

Arxiv

0+阅读 · 2021年9月9日

Sharp Lower Bounds on the Approximation Rate of Shallow Neural Networks

Arxiv

0+阅读 · 2021年9月8日

Wave-Informed Matrix Factorization with Global Optimality Guarantees

Arxiv

0+阅读 · 2021年9月8日

Scaling Properties of Deep Residual Networks

Arxiv

13+阅读 · 2021年5月25日

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Arxiv

13+阅读 · 2020年6月24日

Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Arxiv

8+阅读 · 2018年11月21日

Reducing Parameter Space for Neural Network Training

Arxiv

3+阅读 · 2018年8月17日

Classification with Fairness Constraints: A Meta-Algorithm with Provable Guarantees

Classification with Fairness Constraints: A Meta-Algorithm with Provable Guarantees

Arxiv

3+阅读 · 2018年8月2日

微信扫码咨询专知VIP会员