确定性 MDP 政策循环运行时的一些上层界圈 (Some Upper Bounds on the Running Time of Policy Iteration on Deterministic MDPs) - 专知论文

会员服务 ·

0

策略迭代 · Markov · 相互独立的 · Analysis · Pair ·

2022 年 11 月 28 日

Some Upper Bounds on the Running Time of Policy Iteration on Deterministic MDPs

翻译：确定性 MDP 政策循环运行时的一些上层界圈

Ritesh Goenka,Eashan Gupta,Sushil Khyalia,Pratyush Agarwal,Mulinti Shaik Wajid,Shivaram Kalyanakrishnan

Policy Iteration (PI) is a widely used family of algorithms to compute optimal policies for Markov Decision Problems (MDPs). We derive upper bounds on the running time of PI on Deterministic MDPs (DMDPs): the class of MDPs in which every state-action pair has a unique next state. Our results include a non-trivial upper bound that applies to the entire family of PI algorithms, and affirmation that a conjecture regarding Howard's PI on MDPs is true for DMDPs. Our analysis is based on certain graph-theoretic results, which may be of independent interest.

翻译：政策迭代(PI)是用来计算Markov决策问题最佳政策的一套广泛使用的算法。我们从PI关于决定性 MDP(DMDP)的运行时间中获得上限:每对州-对都有独特的下一个状态的MDP类别。我们的结果包括适用于PI算法整个家族的非三进制上限,并确认Howard关于MDP的PI的猜想对DMDP是真实的。我们的分析基于某些图形理论结果,这些结果可能具有独立的兴趣。

0

相关内容

策略迭代

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

深度强化学习策略梯度教程，53页ppt

深度强化学习策略梯度教程，53页ppt

专知会员服务

184+阅读 · 2020年2月1日

【深度学习架构、模型和技巧集合(TensorFlow/PyTorch)】’Deep Learning Models - A collection of various deep learning architectures, models, and tips'

【深度学习架构、模型和技巧集合(TensorFlow/PyTorch)】’Deep Learning Models - A collection of various deep learning architectures, models, and tips'

专知会员服务

58+阅读 · 2020年1月25日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

SOX9介导TGF-β/Smads/wnt和β-catenin信号通路调控青少年椎体骺板软骨的分化

国家自然科学基金

0+阅读 · 2014年12月31日

miRNAs在XIAP介导的乳腺癌细胞自噬和凋亡平衡性调控中作用及其机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

IL-32/Integrins/FAK通路在肝纤维化形成中的作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

BER通路基因miRNA结合位点基因多态性与结直肠癌易感性的关联及功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

TGF-beta/Smads通路中RBM4参与的转录和转录后水平的调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

lncRNA BC098786和BC088259调控坐骨神经再生的机制

国家自然科学基金

0+阅读 · 2012年12月31日

组蛋白去乙酰化酶抑制剂对骨关节炎中Notch-NFAT信号通路调控的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

黄土丘陵区切沟侵蚀过程的三维数值模拟研究

国家自然科学基金

0+阅读 · 2011年12月31日

hTERT调控相关miRNA的鉴定及功能研究

国家自然科学基金

0+阅读 · 2009年12月31日

条件化恐惧消除过程中海马神经可塑性研究

国家自然科学基金

0+阅读 · 2009年12月31日

Fast Computation of Optimal Transport via Entropy-Regularized Extragradient Methods

Arxiv

0+阅读 · 2023年1月30日

Theoretical analysis of edit distance algorithms: an applied perspective

Arxiv

0+阅读 · 2023年1月30日

Zero-Shot Transfer of Haptics-Based Object Insertion Policies

Arxiv

0+阅读 · 2023年1月29日

All-speed numerical methods for the Euler equations via a sequential explicit time integration

Arxiv

0+阅读 · 2023年1月29日

Efficient Enumeration of Markov Equivalent DAGs

Arxiv

0+阅读 · 2023年1月28日

Learning Mixtures of Markov Chains and MDPs

Arxiv

0+阅读 · 2023年1月28日

An Analysis of Loss Functions for Binary Classification and Regression

Arxiv

0+阅读 · 2023年1月27日

The Stochastic Proximal Distance Algorithm

Arxiv

0+阅读 · 2023年1月27日

f-divergences and their applications in lossy compression and bounding generalization error

Arxiv

0+阅读 · 2023年1月26日

Cluster-Based Control of Transition-Independent MDPs

Arxiv

0+阅读 · 2023年1月26日

VIP会员

文章信息

相关主题

相互独立的

相关VIP内容

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

深度强化学习策略梯度教程，53页ppt

深度强化学习策略梯度教程，53页ppt

专知会员服务

184+阅读 · 2020年2月1日

【深度学习架构、模型和技巧集合(TensorFlow/PyTorch)】’Deep Learning Models - A collection of various deep learning architectures, models, and tips'

【深度学习架构、模型和技巧集合(TensorFlow/PyTorch)】’Deep Learning Models - A collection of various deep learning architectures, models, and tips'

专知会员服务

58+阅读 · 2020年1月25日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军徒步机动作战条令手册》最新168页

【博士论文】基于不确定性的可靠性：现代机器学习中的选择性预测与可信部署

军事后勤数字化未来展望

《美海军后勤体系整合与创新挑战》最新报告

相关资讯

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Fast Computation of Optimal Transport via Entropy-Regularized Extragradient Methods

Arxiv

0+阅读 · 2023年1月30日

Theoretical analysis of edit distance algorithms: an applied perspective

Arxiv

0+阅读 · 2023年1月30日

Zero-Shot Transfer of Haptics-Based Object Insertion Policies

Arxiv

0+阅读 · 2023年1月29日

All-speed numerical methods for the Euler equations via a sequential explicit time integration

Arxiv

0+阅读 · 2023年1月29日

Efficient Enumeration of Markov Equivalent DAGs

Arxiv

0+阅读 · 2023年1月28日

Learning Mixtures of Markov Chains and MDPs

Arxiv

0+阅读 · 2023年1月28日

An Analysis of Loss Functions for Binary Classification and Regression

Arxiv

0+阅读 · 2023年1月27日

The Stochastic Proximal Distance Algorithm

Arxiv

0+阅读 · 2023年1月27日

f-divergences and their applications in lossy compression and bounding generalization error

Arxiv

0+阅读 · 2023年1月26日

Cluster-Based Control of Transition-Independent MDPs

Arxiv

0+阅读 · 2023年1月26日

相关基金

SOX9介导TGF-β/Smads/wnt和β-catenin信号通路调控青少年椎体骺板软骨的分化

国家自然科学基金

0+阅读 · 2014年12月31日

miRNAs在XIAP介导的乳腺癌细胞自噬和凋亡平衡性调控中作用及其机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

IL-32/Integrins/FAK通路在肝纤维化形成中的作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

BER通路基因miRNA结合位点基因多态性与结直肠癌易感性的关联及功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

TGF-beta/Smads通路中RBM4参与的转录和转录后水平的调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

lncRNA BC098786和BC088259调控坐骨神经再生的机制

国家自然科学基金

0+阅读 · 2012年12月31日

组蛋白去乙酰化酶抑制剂对骨关节炎中Notch-NFAT信号通路调控的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

黄土丘陵区切沟侵蚀过程的三维数值模拟研究

国家自然科学基金

0+阅读 · 2011年12月31日

hTERT调控相关miRNA的鉴定及功能研究

国家自然科学基金

0+阅读 · 2009年12月31日

条件化恐惧消除过程中海马神经可塑性研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员