软软体政策梯度和神经复制器动态动态与带加封的暗外探索的内装式政策梯度和神经复制器动态 (Interpolating Between Softmax Policy Gradient and Neural Replicator Dynamics with Capped Implicit Exploration) - 专知论文

会员服务 ·

0

CAP · 赌博机/老虎机 · Softmax · 估计/估计量 · motivation ·

2022 年 6 月 4 日

Interpolating Between Softmax Policy Gradient and Neural Replicator Dynamics with Capped Implicit Exploration

翻译：软软体政策梯度和神经复制器动态动态与带加封的暗外探索的内装式政策梯度和神经复制器动态

Dustin Morrill,Esra'a Saleh,Michael Bowling,Amy Greenwald

from arxiv, At Reinforcement Learning and Decision Making 2022, June 2022. 9 pages and 1 figure

Neural replicator dynamics (NeuRD) is an alternative to the foundational softmax policy gradient (SPG) algorithm motivated by online learning and evolutionary game theory. The NeuRD expected update is designed to be nearly identical to that of SPG, however, we show that the Monte Carlo updates differ in a substantial way: the importance correction accounting for a sampled action is nullified in the SPG update, but not in the NeuRD update. Naturally, this causes the NeuRD update to have higher variance than its SPG counterpart. Building on implicit exploration algorithms in the adversarial bandit setting, we introduce capped implicit exploration (CIX) estimates that allow us to construct NeuRD-CIX, which interpolates between this aspect of NeuRD and SPG. We show how CIX estimates can be used in a black-box reduction to construct bandit algorithms with regret bounds that hold with high probability and the benefits this entails for NeuRD-CIX in sequential decision-making settings. Our analysis reveals a bias--variance tradeoff between SPG and NeuRD, and shows how theory predicts that NeuRD-CIX will perform well more consistently than NeuRD while retaining NeuRD's advantages over SPG in non-stationary environments.

翻译：NeuRD的预期更新设计与SPG几乎完全相同,然而,我们显示蒙特卡洛的更新有很大的不同:抽样行动的重要校正核算在SPG的更新中是无效的,但在NeuRD的更新中则不是。自然,这导致NeurRD的更新与SPG的对应的更新有更大的差异。基于在对抗性强盗设置中的隐性探索算法,我们引入了允许我们建造NeuRD-CIX的上限隐含探索算法(CIX)估计,这在NeuRD和SPG的这一侧面之间是相互对接的。我们展示了CIX的估计数如何在黑盒缩减中使用,以构建具有高度概率的遗憾界限的波段算法,以及由此给NeuRD-CIX在顺序决策环境中带来的好处。我们的分析揭示了SPG和NeurdRD之间的偏差差异,并显示理论将如何预测NeurD-CIX在NEVRD的NURPA环境上比NURD的优势持续保持。

0

相关内容

CAP

CAP原则又称CAP定理，指的是在一个分布式系统中，Consistency（一致性）、 Availability（可用性）、Partition tolerance（分区容错性），三者不可得兼。

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

难治性精神分裂症及其MECT治疗的脑网络特征研究

国家自然科学基金

0+阅读 · 2014年12月31日

氢化对硅烯及硅烯纳米条带电子输运性质的调制

国家自然科学基金

0+阅读 · 2013年12月31日

星形胶质细胞内源性PLD正性调控树突的发育

国家自然科学基金

0+阅读 · 2013年12月31日

采用pinball loss的MEE算法研究

国家自然科学基金

1+阅读 · 2013年12月31日

约束集值优化问题的适定性研究及相关分析

国家自然科学基金

0+阅读 · 2013年12月31日

蛋白激酶GsCBRLK在大豆盐胁迫信号转导途径中的调控机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

Partial Spread Bent函数与Bent-Negabent函数的构造及密码学性质研究

国家自然科学基金

0+阅读 · 2013年12月31日

胶质瘤c-myc/miR-27a靶向SFRP1调控Wnt/β-catenin通路新机制

国家自然科学基金

0+阅读 · 2013年12月31日

胶质母细胞瘤干性起源的分子生物学研究

国家自然科学基金

0+阅读 · 2012年12月31日

胰腺星形细胞对胰腺癌化疗耐药的影响及其机制的研究

国家自然科学基金

0+阅读 · 2011年12月31日

Stronger Generalization Guarantees for Robot Learning by Combining Generative Models and Real-World Data

Stronger Generalization Guarantees for Robot Learning by Combining Generative Models and Real-World Data

Arxiv

0+阅读 · 2022年7月22日

On the sample complexity of stabilizing linear dynamical systems from data

Arxiv

0+阅读 · 2022年7月22日

Reinforcement Learning Approaches for the Orienteering Problem with Stochastic and Dynamic Release Dates

Arxiv

0+阅读 · 2022年7月22日

Robust Knowledge Adaptation for Dynamic Graph Neural Networks

Arxiv

0+阅读 · 2022年7月22日

Differential Geometry for Neural Implicit Models

Arxiv

0+阅读 · 2022年7月21日

Contrastive Learning with Complex Heterogeneity

Contrastive Learning with Complex Heterogeneity

Arxiv

0+阅读 · 2022年7月21日

Learning to Split for Automatic Bias Detection

Arxiv

0+阅读 · 2022年7月20日

The Dice loss in the context of missing or empty labels: Introducing $Φ$ and $ε$

Arxiv

0+阅读 · 2022年7月19日

The Principles of Deep Learning Theory

Arxiv

65+阅读 · 2021年6月18日

Learning with Interpretable Structure from RNN

Arxiv

19+阅读 · 2018年10月25日

VIP会员

文章信息

相关主题

赌博机/老虎机

估计/估计量

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《乌克兰无人机产业：志愿者与政策在构建新兴无人机产业中的协同作用》最新报告

《人工智能辅助决策中的数据可视化：系统性综述》

人工智能驱动弹药制造现代化：美国陆军转型之路

《敏捷作战部署中枢纽-辐条基地选址优化研究》80页

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

相关论文

Stronger Generalization Guarantees for Robot Learning by Combining Generative Models and Real-World Data

Stronger Generalization Guarantees for Robot Learning by Combining Generative Models and Real-World Data

Arxiv

0+阅读 · 2022年7月22日

On the sample complexity of stabilizing linear dynamical systems from data

Arxiv

0+阅读 · 2022年7月22日

Reinforcement Learning Approaches for the Orienteering Problem with Stochastic and Dynamic Release Dates

Arxiv

0+阅读 · 2022年7月22日

Robust Knowledge Adaptation for Dynamic Graph Neural Networks

Arxiv

0+阅读 · 2022年7月22日

Differential Geometry for Neural Implicit Models

Arxiv

0+阅读 · 2022年7月21日

Contrastive Learning with Complex Heterogeneity

Contrastive Learning with Complex Heterogeneity

Arxiv

0+阅读 · 2022年7月21日

Learning to Split for Automatic Bias Detection

Arxiv

0+阅读 · 2022年7月20日

The Dice loss in the context of missing or empty labels: Introducing $Φ$ and $ε$

Arxiv

0+阅读 · 2022年7月19日

The Principles of Deep Learning Theory

Arxiv

65+阅读 · 2021年6月18日

Learning with Interpretable Structure from RNN

Arxiv

19+阅读 · 2018年10月25日

相关基金

难治性精神分裂症及其MECT治疗的脑网络特征研究

国家自然科学基金

0+阅读 · 2014年12月31日

氢化对硅烯及硅烯纳米条带电子输运性质的调制

国家自然科学基金

0+阅读 · 2013年12月31日

星形胶质细胞内源性PLD正性调控树突的发育

国家自然科学基金

0+阅读 · 2013年12月31日

采用pinball loss的MEE算法研究

国家自然科学基金

1+阅读 · 2013年12月31日

约束集值优化问题的适定性研究及相关分析

国家自然科学基金

0+阅读 · 2013年12月31日

蛋白激酶GsCBRLK在大豆盐胁迫信号转导途径中的调控机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

Partial Spread Bent函数与Bent-Negabent函数的构造及密码学性质研究

国家自然科学基金

0+阅读 · 2013年12月31日

胶质瘤c-myc/miR-27a靶向SFRP1调控Wnt/β-catenin通路新机制

国家自然科学基金

0+阅读 · 2013年12月31日

胶质母细胞瘤干性起源的分子生物学研究

国家自然科学基金

0+阅读 · 2012年12月31日

胰腺星形细胞对胰腺癌化疗耐药的影响及其机制的研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员