通过多层次蒙特卡洛行为者-批评,在平均奖励强化学习中超越指数快速混合 (Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-Critic) - 专知论文

会员服务 ·

0

混合时间 · 蒙特卡罗 · 混合 · FAST · Learning ·

2023 年 2 月 1 日

Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-Critic

翻译：通过多层次蒙特卡洛行为者-批评,在平均奖励强化学习中超越指数快速混合

Wesley A. Suttle,Amrit Singh Bedi,Bhrij Patel,Brian M. Sadler,Alec Koppel,Dinesh Manocha

Many existing reinforcement learning (RL) methods employ stochastic gradient iteration on the back end, whose stability hinges upon a hypothesis that the data-generating process mixes exponentially fast with a rate parameter that appears in the step-size selection. Unfortunately, this assumption is violated for large state spaces or settings with sparse rewards, and the mixing time is unknown, making the step size inoperable. In this work, we propose an RL methodology attuned to the mixing time by employing a multi-level Monte Carlo estimator for the critic, the actor, and the average reward embedded within an actor-critic (AC) algorithm. This method, which we call \textbf{M}ulti-level \textbf{A}ctor-\textbf{C}ritic (MAC), is developed especially for infinite-horizon average-reward settings and neither relies on oracle knowledge of the mixing time in its parameter selection nor assumes its exponential decay; it, therefore, is readily applicable to applications with slower mixing times. Nonetheless, it achieves a convergence rate comparable to the state-of-the-art AC algorithms. We experimentally show that these alleviated restrictions on the technical conditions required for stability translate to superior performance in practice for RL problems with sparse rewards.

翻译：许多现有的强化学习(RL)方法在后端采用随机梯度变异,其稳定性取决于以下假设:数据生成过程与在步数选择中出现的速率参数成倍地快速混合。不幸的是,对于大型国家空间或环境而言,这一假设被违反,回报微弱,混合时间不详,使步数无法操作。在这项工作中,我们提议一种RL方法,通过对评论家、演员和行为者采用多层次的蒙特卡洛测深器来适应混合时间,以及行为者-捷克算法(AC)中包含的平均奖赏。尽管如此,我们称之为\textbf{M}multi-le level level \ textbf{A}}A}ctor- textbf{C}ritic (MACC) 的方法,特别为无限偏差平均反向环境开发了这一假设,既不依赖其参数选择中混合时间的知识,也不假设其指数衰减;因此,该方法很容易适用于混合时间较慢的应用。尽管如此,它还是实现了与州-州-艺术混合算算算出的趋近的趋近的趋近的趋近的趋同AC级演算法,但我们用这些技术测测测测测测测测得这些技术问题。

0

相关内容

混合时间

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

溶解性胶粉改性沥青的微细观结构与流变性能研究

国家自然科学基金

0+阅读 · 2014年12月31日

中低温固体氧化物燃料电池核-壳结构阴极的研究

国家自然科学基金

0+阅读 · 2014年12月31日

固体氧化物燃料电池纳米结构阴极的构筑及中低温电化学性能

国家自然科学基金

0+阅读 · 2014年12月31日

压水堆PCI风险控制策略研究

国家自然科学基金

1+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

Kronheimer-Nakajima quiver 模空间与有理曲面

国家自然科学基金

1+阅读 · 2013年12月31日

纳米结构SOFC复合阴极的动力学过程研究

国家自然科学基金

0+阅读 · 2013年12月31日

复杂有机膦酸盐的结构与性能

国家自然科学基金

0+阅读 · 2012年12月31日

电沉积氧化石墨烯/ZnO-SnO2纳米复合膜的光电转换性能

国家自然科学基金

0+阅读 · 2011年12月31日

p进表示的伽罗瓦上同调

国家自然科学基金

0+阅读 · 2008年12月31日

ReBotNet: Fast Real-time Video Enhancement

ReBotNet: Fast Real-time Video Enhancement

Arxiv

0+阅读 · 2023年3月23日

Sample-Efficient Multi-Objective Learning via Generalized Policy Improvement Prioritization

Arxiv

0+阅读 · 2023年3月23日

Semi-Oblivious Chase Termination for Linear Existential Rules: An Experimental Study

Arxiv

0+阅读 · 2023年3月22日

Guiding Online Reinforcement Learning with Action-Free Offline Pretraining

Arxiv

0+阅读 · 2023年3月22日

Learning Stationary Nash Equilibrium Policies in $n$-Player Stochastic Games with Independent Chains

Arxiv

0+阅读 · 2023年3月22日

Stateless actor-critic for instance segmentation with high-level priors

Stateless actor-critic for instance segmentation with high-level priors

Arxiv

0+阅读 · 2023年3月21日

Bandits Corrupted by Nature: Lower Bounds on Regret and Robust Optimistic Algorithm

Arxiv

0+阅读 · 2023年3月21日

Multi-Resolution Online Deterministic Annealing: A Hierarchical and Progressive Learning Architecture

Arxiv

0+阅读 · 2023年3月21日

Fast exploration and learning of latent graphs with aliased observations

Arxiv

0+阅读 · 2023年3月21日

Bridging Imitation and Online Reinforcement Learning: An Optimistic Tale

Arxiv

0+阅读 · 2023年3月20日

VIP会员

文章信息

相关主题

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《小型无人机系统侦测追踪技术：声学、计算机视觉与深度学习融合方案》最新98页

《"牧羊人网格"拦截策略：实现无人机集群可靠拦截的新范式》

光纤无人机：反无人机系统的重大挑战

《作战建模与仿真实证研究》

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

相关论文

ReBotNet: Fast Real-time Video Enhancement

ReBotNet: Fast Real-time Video Enhancement

Arxiv

0+阅读 · 2023年3月23日

Sample-Efficient Multi-Objective Learning via Generalized Policy Improvement Prioritization

Arxiv

0+阅读 · 2023年3月23日

Semi-Oblivious Chase Termination for Linear Existential Rules: An Experimental Study

Arxiv

0+阅读 · 2023年3月22日

Guiding Online Reinforcement Learning with Action-Free Offline Pretraining

Arxiv

0+阅读 · 2023年3月22日

Learning Stationary Nash Equilibrium Policies in $n$-Player Stochastic Games with Independent Chains

Arxiv

0+阅读 · 2023年3月22日

Stateless actor-critic for instance segmentation with high-level priors

Stateless actor-critic for instance segmentation with high-level priors

Arxiv

0+阅读 · 2023年3月21日

Bandits Corrupted by Nature: Lower Bounds on Regret and Robust Optimistic Algorithm

Arxiv

0+阅读 · 2023年3月21日

Multi-Resolution Online Deterministic Annealing: A Hierarchical and Progressive Learning Architecture

Arxiv

0+阅读 · 2023年3月21日

Fast exploration and learning of latent graphs with aliased observations

Arxiv

0+阅读 · 2023年3月21日

Bridging Imitation and Online Reinforcement Learning: An Optimistic Tale

Arxiv

0+阅读 · 2023年3月20日

相关基金

溶解性胶粉改性沥青的微细观结构与流变性能研究

国家自然科学基金

0+阅读 · 2014年12月31日

中低温固体氧化物燃料电池核-壳结构阴极的研究

国家自然科学基金

0+阅读 · 2014年12月31日

固体氧化物燃料电池纳米结构阴极的构筑及中低温电化学性能

国家自然科学基金

0+阅读 · 2014年12月31日

压水堆PCI风险控制策略研究

国家自然科学基金

1+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

Kronheimer-Nakajima quiver 模空间与有理曲面

国家自然科学基金

1+阅读 · 2013年12月31日

纳米结构SOFC复合阴极的动力学过程研究

国家自然科学基金

0+阅读 · 2013年12月31日

复杂有机膦酸盐的结构与性能

国家自然科学基金

0+阅读 · 2012年12月31日

电沉积氧化石墨烯/ZnO-SnO2纳米复合膜的光电转换性能

国家自然科学基金

0+阅读 · 2011年12月31日

p进表示的伽罗瓦上同调

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员