通过更多探索进行动态零散培训 (Dynamic Sparse Training via More Exploration) - 专知论文

会员服务 ·

0

稀疏 · 模型评估 · MoDELS · 泛函 · 可约的 ·

2022 年 12 月 14 日

Dynamic Sparse Training via More Exploration

翻译：通过更多探索进行动态零散培训

Shaoyi Huang,Bowen Lei,Dongkuan Xu,Hongwu Peng,Yue Sun,Mimi Xie,Caiwen Ding

Over-parameterization of deep neural networks (DNNs) has shown high prediction accuracy for many applications. Although effective, the large number of parameters hinders its popularity on resource-limited devices and has an outsize environmental impact. Sparse training (using a fixed number of nonzero weights in each iteration) could significantly mitigate the training costs by reducing the model size. However, existing sparse training methods mainly use either random-based or greedy-based drop-and-grow strategies, resulting in local minimal and low accuracy. In this work, we consider the dynamic sparse training as a sparse connectivity search problem and design an exploitation and exploration acquisition function to escape from local optima and saddle points. We further design an acquisition function and provide the theoretical guarantees for the proposed method and clarify its convergence property. Experimental results show that sparse models (up to 98\% sparsity) obtained by our proposed method outperform the SOTA sparse training methods on a wide variety of deep learning tasks. On VGG-19 / CIFAR-100, ResNet-50 / CIFAR-10, ResNet-50 / CIFAR-100, our method has even higher accuracy than dense models. On ResNet-50 / ImageNet, the proposed method has up to 8.2\% accuracy improvement compared to SOTA sparse training methods.

翻译：深度神经网络(DNNs)的超度测量显示,许多应用的预测准确度很高,尽管效果有效,但大量参数的众多参数妨碍了其受资源有限的装置的欢迎程度,并对环境产生了超大的影响。粗糙的培训(在每迭中使用固定数量的非零重量)可以通过缩小模型规模,大大降低培训成本。然而,现有的稀少的培训方法主要使用随机或贪婪的滴滴滴式战略,导致当地最低和低精确度。在这项工作中,我们认为动态的稀少培训是一个稀少的连通搜索问题,设计了一个探索和勘探获取功能,以逃避当地Opima和垫接点。我们进一步设计了一个获取功能,为拟议方法提供理论保障,并澄清其趋同性。实验结果表明,我们拟议方法获得的稀少模式(高达98 ⁇ 宽度)超过了SOTA的分散培训方法,在各种深层次的学习任务方面,使SOTA少的训练方法更加完美。关于VGG-19/CIFAR-100,ResNet-50/CIFAR-100,我们的方法网-50/CIFAR-100,我们的方法比ROFAR-AR-100,我们的拟议方法的精确性方法比QAR-RO50/SORION-AR-LOLODAS-IAR-IAR-LA模型的精确度也比高。

0

相关内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【SIGIR2018】五篇对抗训练文章

【SIGIR2018】五篇对抗训练文章

专知

12+阅读 · 2018年7月9日

【推荐】深度学习目标检测概览

【推荐】深度学习目标检测概览

机器学习研究会

10+阅读 · 2017年9月1日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

基于海量软件片段比对的恶意代码检测方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

樟疫霉致病性相关GPCR-PIPK鉴定与机理研究

国家自然科学基金

0+阅读 · 2015年12月31日

领域驱动空间co-location模式挖掘技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

纳米金属微观结构不稳定性机理的三维研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于宏基因组文库的新壳聚糖酶筛选及计算机辅助合理化设计

国家自然科学基金

0+阅读 · 2013年12月31日

肿瘤标志物的单分子检测与成像分析方法的研究

国家自然科学基金

0+阅读 · 2012年12月31日

热障涂层的冲蚀破坏机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

TiNi形状记忆合金表面W离子注入改性及其机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

DEC1、DEC2对人乳腺癌细胞衰老的调控作用及其作用机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Skutterudite/AgSbTe2系纳米复合热电材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

Understanding Expertise through Demonstrations: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning

Arxiv

0+阅读 · 2023年2月15日

Adaptive design of experiment via normalizing flows for failure probability estimation

Arxiv

0+阅读 · 2023年2月14日

On the Computational Efficiency of Adaptive and Dynamic Regret Minimization

On the Computational Efficiency of Adaptive and Dynamic Regret Minimization

Arxiv

0+阅读 · 2023年2月13日

Automatic Noise Filtering with Dynamic Sparse Training in Deep Reinforcement Learning

Arxiv

0+阅读 · 2023年2月13日

Density-Softmax: Scalable and Distance-Aware Uncertainty Estimation under Distribution Shifts

Arxiv

0+阅读 · 2023年2月13日

Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex Optimization

Arxiv

0+阅读 · 2023年2月12日

Approximation and Structured Prediction with Sparse Wasserstein Barycenters

Arxiv

0+阅读 · 2023年2月10日

More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization

Arxiv

0+阅读 · 2023年2月10日

Active Learning for Domain Adaptation: An Energy-based Approach

Arxiv

13+阅读 · 2021年12月2日

On Explainability of Graph Neural Networks via Subgraph Explorations

Arxiv

11+阅读 · 2021年5月31日

VIP会员

文章信息

相关主题

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

扩散语言模型综述

《美陆军徒步机动作战条令手册》最新168页

【博士论文】理解神经网络的训练动态：从局部优化轨迹与特征学习视角

军事后勤数字化未来展望

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【SIGIR2018】五篇对抗训练文章

【SIGIR2018】五篇对抗训练文章

专知

12+阅读 · 2018年7月9日

【推荐】深度学习目标检测概览

【推荐】深度学习目标检测概览

机器学习研究会

10+阅读 · 2017年9月1日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Understanding Expertise through Demonstrations: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning

Arxiv

0+阅读 · 2023年2月15日

Adaptive design of experiment via normalizing flows for failure probability estimation

Arxiv

0+阅读 · 2023年2月14日

On the Computational Efficiency of Adaptive and Dynamic Regret Minimization

On the Computational Efficiency of Adaptive and Dynamic Regret Minimization

Arxiv

0+阅读 · 2023年2月13日

Automatic Noise Filtering with Dynamic Sparse Training in Deep Reinforcement Learning

Arxiv

0+阅读 · 2023年2月13日

Density-Softmax: Scalable and Distance-Aware Uncertainty Estimation under Distribution Shifts

Arxiv

0+阅读 · 2023年2月13日

Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex Optimization

Arxiv

0+阅读 · 2023年2月12日

Approximation and Structured Prediction with Sparse Wasserstein Barycenters

Arxiv

0+阅读 · 2023年2月10日

More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization

Arxiv

0+阅读 · 2023年2月10日

Active Learning for Domain Adaptation: An Energy-based Approach

Arxiv

13+阅读 · 2021年12月2日

On Explainability of Graph Neural Networks via Subgraph Explorations

Arxiv

11+阅读 · 2021年5月31日

相关基金

基于海量软件片段比对的恶意代码检测方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

樟疫霉致病性相关GPCR-PIPK鉴定与机理研究

国家自然科学基金

0+阅读 · 2015年12月31日

领域驱动空间co-location模式挖掘技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

纳米金属微观结构不稳定性机理的三维研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于宏基因组文库的新壳聚糖酶筛选及计算机辅助合理化设计

国家自然科学基金

0+阅读 · 2013年12月31日

肿瘤标志物的单分子检测与成像分析方法的研究

国家自然科学基金

0+阅读 · 2012年12月31日

热障涂层的冲蚀破坏机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

TiNi形状记忆合金表面W离子注入改性及其机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

DEC1、DEC2对人乳腺癌细胞衰老的调控作用及其作用机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Skutterudite/AgSbTe2系纳米复合热电材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员