政策分级方法对综合国家产生的近似效益 (Approximation Benefits of Policy Gradient Methods with Aggregated States) - 专知论文

会员服务 ·

0

近似 · 价值函数 · 稳健性 · 近似误差 · 值函数近似 ·

2022 年 6 月 23 日

Approximation Benefits of Policy Gradient Methods with Aggregated States

翻译：政策分级方法对综合国家产生的近似效益

Folklore suggests that policy gradient can be more robust to misspecification than its relative, approximate policy iteration. This paper studies the case of state-aggregated representations, where the state space is partitioned and either the policy or value function approximation is held constant over partitions. This paper shows a policy gradient method converges to a policy whose regret per-period is bounded by $\epsilon$, the largest difference between two elements of the state-action value function belonging to a common partition. With the same representation, both approximate policy iteration and approximate value iteration can produce policies whose per-period regret scales as $\epsilon/(1-\gamma)$, where $\gamma$ is a discount factor. Faced with inherent approximation error, methods that locally optimize the true decision-objective can be far more robust.

翻译：民俗认为,政策梯度比其相对的、近似的政策迭代更强, 比其相对的、近似的政策迭代更强。本文研究了国家汇总的表述方式, 国家空间被分割, 政策或价值函数近似在分割区上保持恒定。本文显示了一种政策梯度方法, 政策梯度方法与一个政策相趋一致, 该政策对每期的遗憾是$- epsilon 美元, 这是属于共同分割区的国家- 行动价值函数中两个元素的最大差异。以同样的表述方式, 政策迭代和近似值迭代可以产生每期的遗憾等级为$\ epsilon/ (1-\ gamma)$( 1-\ gamma) 的政策, 美元是一个折扣系数。面对内在的近似错误, 当地优化真正决策目标的方法可能更加稳健。

0

相关内容

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

不确定分数阶非线性系统Mittag-Leffler自适应控制

国家自然科学基金

1+阅读 · 2016年12月31日

S3AGA样本（Spitzer-SDSS Spectral Atlas of Galaxies and AGNs)及其AGN研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于高电子迁移率电极材料的高效钙钛矿太阳电池的研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于轨迹灵敏度的随机网络控制系统不敏感控制

国家自然科学基金

1+阅读 · 2013年12月31日

基于谓词抽象技术的访问控制策略安全性快速判定方法的研究

国家自然科学基金

1+阅读 · 2013年12月31日

全参数完美匹配层的人工实现

国家自然科学基金

0+阅读 · 2013年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

基于水溶性共轭聚合物的DNA序列分析研究及其在DNA微阵列技术中的应用

国家自然科学基金

0+阅读 · 2009年12月31日

超声组装掺杂半导体纳米材料与电致化学发光生物传感

国家自然科学基金

0+阅读 · 2009年12月31日

用于同位素18O光纤低损耗窗口（1730-1760nm）增益平坦的石英基Tm:Ho共掺光纤放大器研制

国家自然科学基金

0+阅读 · 2008年12月31日

Backward error analysis for conjugate symplectic methods

Arxiv

0+阅读 · 2022年8月12日

Improving performance in multi-objective decision-making in Bottles environments with soft maximin approaches

Arxiv

0+阅读 · 2022年8月11日

Off-Policy Actor-Critic with Emphatic Weightings

Arxiv

0+阅读 · 2022年8月11日

Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Streaming Data

Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Streaming Data

Arxiv

0+阅读 · 2022年8月11日

Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity

Arxiv

0+阅读 · 2022年8月11日

Adaptive Learning Rates for Faster Stochastic Gradient Methods

Arxiv

0+阅读 · 2022年8月10日

Arbitrary High Order WENO Finite Volume Scheme with Flux Globalization for Moving Equilibria Preservation

Arxiv

0+阅读 · 2022年8月10日

Deep Learning Methods for Proximal Inference via Maximum Moment Restriction

Arxiv

0+阅读 · 2022年8月9日

Decomposed Mutual Information Estimation for Contrastive Representation Learning

Arxiv

11+阅读 · 2021年6月25日

IEOPF: An Active Contour Model for Image Segmentation with Inhomogeneities Estimated by Orthogonal Primary Functions

Arxiv

10+阅读 · 2018年1月20日

VIP会员

文章信息

相关主题

值函数近似

相关VIP内容

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津博士论文】零样本强化学习综述

《美军条令：陆军指挥官与规划人员地理空间指南》60页

战术边缘指挥控制：防务面临的核心挑战

迈向开放世界检测：综述

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

相关论文

Backward error analysis for conjugate symplectic methods

Arxiv

0+阅读 · 2022年8月12日

Improving performance in multi-objective decision-making in Bottles environments with soft maximin approaches

Arxiv

0+阅读 · 2022年8月11日

Off-Policy Actor-Critic with Emphatic Weightings

Arxiv

0+阅读 · 2022年8月11日

Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Streaming Data

Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Streaming Data

Arxiv

0+阅读 · 2022年8月11日

Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity

Arxiv

0+阅读 · 2022年8月11日

Adaptive Learning Rates for Faster Stochastic Gradient Methods

Arxiv

0+阅读 · 2022年8月10日

Arbitrary High Order WENO Finite Volume Scheme with Flux Globalization for Moving Equilibria Preservation

Arxiv

0+阅读 · 2022年8月10日

Deep Learning Methods for Proximal Inference via Maximum Moment Restriction

Arxiv

0+阅读 · 2022年8月9日

Decomposed Mutual Information Estimation for Contrastive Representation Learning

Arxiv

11+阅读 · 2021年6月25日

IEOPF: An Active Contour Model for Image Segmentation with Inhomogeneities Estimated by Orthogonal Primary Functions

Arxiv

10+阅读 · 2018年1月20日

相关基金

不确定分数阶非线性系统Mittag-Leffler自适应控制

国家自然科学基金

1+阅读 · 2016年12月31日

S3AGA样本（Spitzer-SDSS Spectral Atlas of Galaxies and AGNs)及其AGN研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于高电子迁移率电极材料的高效钙钛矿太阳电池的研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于轨迹灵敏度的随机网络控制系统不敏感控制

国家自然科学基金

1+阅读 · 2013年12月31日

基于谓词抽象技术的访问控制策略安全性快速判定方法的研究

国家自然科学基金

1+阅读 · 2013年12月31日

全参数完美匹配层的人工实现

国家自然科学基金

0+阅读 · 2013年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

基于水溶性共轭聚合物的DNA序列分析研究及其在DNA微阵列技术中的应用

国家自然科学基金

0+阅读 · 2009年12月31日

超声组装掺杂半导体纳米材料与电致化学发光生物传感

国家自然科学基金

0+阅读 · 2009年12月31日

用于同位素18O光纤低损耗窗口（1730-1760nm）增益平坦的石英基Tm:Ho共掺光纤放大器研制

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员