一般政策优化分析更新规则 (An Analytical Update Rule for General Policy Optimization) - 专知论文

会员服务 ·

0

随机性策略 · 优化器 · 策略搜索 · Performer · 讲稿 ·

2022 年 5 月 14 日

An Analytical Update Rule for General Policy Optimization

翻译：一般政策优化分析更新规则

Hepeng Li,Nicholas Clavette,Haibo He

We present an analytical policy update rule that is independent of parameterized function approximators. The update rule is suitable for general stochastic policies with monotonic improvement guarantee. The update rule is derived from a closed-form trust-region solution using calculus of variation, following a new theoretical result that tightens existing bounds for policy search using trust-region methods. An explanation building a connection between the policy update rule and value-function methods is provided. Based on a recursive form of the update rule, an off-policy algorithm is derived naturally, and the monotonic improvement guarantee remains. Furthermore, the update rule extends immediately to multi-agent systems when updates are performed by one agent at a time.

翻译：我们提出了一个独立于参数化功能近似器的分析性政策更新规则。更新规则适合于具有单声道改进保证的一般随机政策。更新规则源于一种使用变异计算法的封闭式信任区域解决办法,该计算法采用新的理论结果,用信任区域方法收紧了政策搜索的现有界限。提供了在政策更新规则和价值功能方法之间建立联系的解释。根据更新规则的循环形式,自然产生一种非政策性算法,单声道改进保证仍然存在。此外,更新规则在由一个代理人一次进行更新时立即扩展到多剂系统。

0

相关内容

随机性策略

随机性策略

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

59+阅读 · 2022年4月22日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

50+阅读 · 2020年12月14日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

70+阅读 · 2020年8月2日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

75+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

91+阅读 · 2020年3月12日

深度强化学习策略梯度教程，53页ppt

深度强化学习策略梯度教程，53页ppt

专知会员服务

176+阅读 · 2020年2月1日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

开源书：PyTorch深度学习起步

开源书：PyTorch深度学习起步

专知会员服务

49+阅读 · 2019年10月11日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

167+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

90+阅读 · 2019年10月10日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

23+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

25+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

41+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

16+阅读 · 2018年12月24日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

小麦TaNPF8.1(B)基因在氮素吸收利用中的功能研究

国家自然科学基金

0+阅读 · 2015年12月31日

HIF-1/COMPASS调控缺氧诱导Brg1和Brm表达上调的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

miRNAs调控柿单宁合成代谢机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

利用钙离子魔幻波长高精度测量4S-4P跃迁振子强度

国家自然科学基金

0+阅读 · 2014年12月31日

多基因协同调控里氏木霉高效分泌表达蛋白的研究

国家自然科学基金

0+阅读 · 2013年12月31日

柽柳Dof转录因子的耐盐调控机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

Kupffer细胞上GITRL在大鼠肝移植免疫耐受重建中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于ForCES的软件定义网络（SDN）研究

国家自然科学基金

1+阅读 · 2012年12月31日

UGT基因簇进化及调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

银耳二型态相关基因的克隆与鉴定

国家自然科学基金

0+阅读 · 2009年12月31日

Families of polytopes with rational linear precision in higher dimensions

Arxiv

0+阅读 · 2022年7月5日

Global Convergence of Successive Approximations for Non-convex Stochastic Optimal Control Problems

Arxiv

0+阅读 · 2022年7月5日

Learning Stochastic Shortest Path with Linear Function Approximation

Arxiv

0+阅读 · 2022年7月5日

A Simple and Optimal Policy Design with Safety against Heavy-tailed Risk for Multi-armed Bandits

Arxiv

0+阅读 · 2022年7月4日

On the Number of Quantifiers as a Complexity Measure

Arxiv

0+阅读 · 2022年7月4日

Optimal Extragradient-Based Bilinearly-Coupled Saddle-Point Optimization

Arxiv

0+阅读 · 2022年7月3日

DeepAL for Regression Using $ε$-weighted Hybrid Query Strategy

Arxiv

0+阅读 · 2022年7月3日

Empirical likelihood inference for longitudinal data with covariate measurement errors: An application to the LEAN study

Arxiv

0+阅读 · 2022年7月2日

Transformers are Meta-Reinforcement Learners

Arxiv

15+阅读 · 2022年6月14日

The Causal Learning of Retail Delinquency

Arxiv

14+阅读 · 2020年12月17日

VIP会员

文章信息

相关主题

随机性策略

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

59+阅读 · 2022年4月22日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

50+阅读 · 2020年12月14日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

70+阅读 · 2020年8月2日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

75+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

91+阅读 · 2020年3月12日

深度强化学习策略梯度教程，53页ppt

深度强化学习策略梯度教程，53页ppt

专知会员服务

176+阅读 · 2020年2月1日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

开源书：PyTorch深度学习起步

开源书：PyTorch深度学习起步

专知会员服务

49+阅读 · 2019年10月11日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

167+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

90+阅读 · 2019年10月10日

热门VIP内容

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

23+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

25+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

41+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

16+阅读 · 2018年12月24日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Families of polytopes with rational linear precision in higher dimensions

Arxiv

0+阅读 · 2022年7月5日

Global Convergence of Successive Approximations for Non-convex Stochastic Optimal Control Problems

Arxiv

0+阅读 · 2022年7月5日

Learning Stochastic Shortest Path with Linear Function Approximation

Arxiv

0+阅读 · 2022年7月5日

A Simple and Optimal Policy Design with Safety against Heavy-tailed Risk for Multi-armed Bandits

Arxiv

0+阅读 · 2022年7月4日

On the Number of Quantifiers as a Complexity Measure

Arxiv

0+阅读 · 2022年7月4日

Optimal Extragradient-Based Bilinearly-Coupled Saddle-Point Optimization

Arxiv

0+阅读 · 2022年7月3日

DeepAL for Regression Using $ε$-weighted Hybrid Query Strategy

Arxiv

0+阅读 · 2022年7月3日

Empirical likelihood inference for longitudinal data with covariate measurement errors: An application to the LEAN study

Arxiv

0+阅读 · 2022年7月2日

Transformers are Meta-Reinforcement Learners

Arxiv

15+阅读 · 2022年6月14日

The Causal Learning of Retail Delinquency

Arxiv

14+阅读 · 2020年12月17日

相关基金

小麦TaNPF8.1(B)基因在氮素吸收利用中的功能研究

国家自然科学基金

0+阅读 · 2015年12月31日

HIF-1/COMPASS调控缺氧诱导Brg1和Brm表达上调的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

miRNAs调控柿单宁合成代谢机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

利用钙离子魔幻波长高精度测量4S-4P跃迁振子强度

国家自然科学基金

0+阅读 · 2014年12月31日

多基因协同调控里氏木霉高效分泌表达蛋白的研究

国家自然科学基金

0+阅读 · 2013年12月31日

柽柳Dof转录因子的耐盐调控机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

Kupffer细胞上GITRL在大鼠肝移植免疫耐受重建中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于ForCES的软件定义网络（SDN）研究

国家自然科学基金

1+阅读 · 2012年12月31日

UGT基因簇进化及调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

银耳二型态相关基因的克隆与鉴定

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员