以依赖样本为基础的Vanilla示范非线外强化学习样本复杂性</s> (On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent Samples) - 专知论文

会员服务 ·

0

样本复杂度 · 样本 · 估计/估计量 · Learning · 情景 ·

2023 年 3 月 7 日

On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent Samples

翻译：以依赖样本为基础的Vanilla示范非线外强化学习样本复杂性

Mustafa O. Karabag,Ufuk Topcu

from arxiv, Accepted to AAAI-23

Offline reinforcement learning (offline RL) considers problems where learning is performed using only previously collected samples and is helpful for the settings in which collecting new data is costly or risky. In model-based offline RL, the learner performs estimation (or optimization) using a model constructed according to the empirical transition frequencies. We analyze the sample complexity of vanilla model-based offline RL with dependent samples in the infinite-horizon discounted-reward setting. In our setting, the samples obey the dynamics of the Markov decision process and, consequently, may have interdependencies. Under no assumption of independent samples, we provide a high-probability, polynomial sample complexity bound for vanilla model-based off-policy evaluation that requires partial or uniform coverage. We extend this result to the off-policy optimization under uniform coverage. As a comparison to the model-based approach, we analyze the sample complexity of off-policy evaluation with vanilla importance sampling in the infinite-horizon setting. Finally, we provide an estimator that outperforms the sample-mean estimator for almost deterministic dynamics that are prevalent in reinforcement learning.

翻译：离线强化学习(离线 RL) 考虑的是,在学习过程中,仅使用先前收集的样本进行学习的问题,对于收集新数据的成本或风险环境是有助益的。在基于模型的离线RL中,学习者使用根据经验过渡频率构建的模式进行估计(或优化)。我们分析了香草模型离线学习的样本复杂性,在无穷度折扣降价环境下有依赖性的样本。在设定中,样本符合Markov决定过程的动态,因此可能具有相互依存性。在不假定独立样本的情况下,我们提供了一种高概率、多元样本复杂性,用于需要部分或统一覆盖的香草模型离政策评估。我们将这一结果推广到统一覆盖下的脱政策优化。作为与基于模型的方法的比较,我们分析了非政策评估的样本复杂性,在无穷度休养分层设置中,我们提供了一种估测算器,它比样本中标定值的精度高,因为几乎具有确定性动态,在加强学习中十分普遍。</s>

0

相关内容

样本复杂度

样本复杂度

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【MIla】一种意识启发规划的基于模型强化学习，A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning

【MIla】一种意识启发规划的基于模型强化学习，A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning

专知会员服务

23+阅读 · 2022年3月19日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

53+阅读 · 2020年9月7日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

【推荐】GAN架构入门综述(资源汇总)

【推荐】GAN架构入门综述(资源汇总)

机器学习研究会

10+阅读 · 2017年9月3日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

采用pinball loss的MEE算法研究

国家自然科学基金

1+阅读 · 2013年12月31日

CTRP1促进紊流下动脉内皮屏障功能损伤及其机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

深共熔溶剂介质中镍磷合金的可控制备及电化学性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

离子液体中纤维素直接转化为5-羟甲基糠醛催化反应研究

国家自然科学基金

0+阅读 · 2012年12月31日

近似计数的算法与复杂性

国家自然科学基金

1+阅读 · 2012年12月31日

天然产物Artanomalide D及其类似物的全合成和抗肿瘤构效关系研究

国家自然科学基金

0+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

吡唑基铜配合物的合成与组装

国家自然科学基金

0+阅读 · 2011年12月31日

基于可达域计算的AWID-AWIS车辆动力学协调控制

国家自然科学基金

0+阅读 · 2010年12月31日

A Generic Approach for Reproducible Model Distillation

Arxiv

0+阅读 · 2023年4月27日

Adversarial Policy Optimization in Deep Reinforcement Learning

Arxiv

0+阅读 · 2023年4月27日

Linear and Nonlinear Parareal Methods for the Cahn-Hilliard Equation

Arxiv

0+阅读 · 2023年4月27日

One-Step Distributional Reinforcement Learning

Arxiv

0+阅读 · 2023年4月27日

Online Spatio-Temporal Learning with Target Projection

Arxiv

0+阅读 · 2023年4月26日

A Survey of Meta-Reinforcement Learning

Arxiv

12+阅读 · 2023年1月19日

Reinforcement Learning on Graph: A Survey

Arxiv

67+阅读 · 2022年4月13日

Causality and Generalizability: Identifiability and Learning Methods

Arxiv

12+阅读 · 2021年10月4日

Coding for Distributed Multi-Agent Reinforcement Learning

Arxiv

32+阅读 · 2021年1月7日

A Multi-Objective Deep Reinforcement Learning Framework

A Multi-Objective Deep Reinforcement Learning Framework

Arxiv

16+阅读 · 2018年6月27日

VIP会员

文章信息

相关主题

样本复杂度

估计/估计量

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【MIla】一种意识启发规划的基于模型强化学习，A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning

【MIla】一种意识启发规划的基于模型强化学习，A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning

专知会员服务

23+阅读 · 2022年3月19日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

53+阅读 · 2020年9月7日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】以人为中心的强化学习

任务规划与地形分析：现代复杂环境作战导航体系

认知优势：人工智能在国家安全决策中的核心作用

大模型赋能的具身智能：决策与具身学习综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

【推荐】GAN架构入门综述(资源汇总)

【推荐】GAN架构入门综述(资源汇总)

机器学习研究会

10+阅读 · 2017年9月3日

相关论文

A Generic Approach for Reproducible Model Distillation

Arxiv

0+阅读 · 2023年4月27日

Adversarial Policy Optimization in Deep Reinforcement Learning

Arxiv

0+阅读 · 2023年4月27日

Linear and Nonlinear Parareal Methods for the Cahn-Hilliard Equation

Arxiv

0+阅读 · 2023年4月27日

One-Step Distributional Reinforcement Learning

Arxiv

0+阅读 · 2023年4月27日

Online Spatio-Temporal Learning with Target Projection

Arxiv

0+阅读 · 2023年4月26日

A Survey of Meta-Reinforcement Learning

Arxiv

12+阅读 · 2023年1月19日

Reinforcement Learning on Graph: A Survey

Arxiv

67+阅读 · 2022年4月13日

Causality and Generalizability: Identifiability and Learning Methods

Arxiv

12+阅读 · 2021年10月4日

Coding for Distributed Multi-Agent Reinforcement Learning

Arxiv

32+阅读 · 2021年1月7日

A Multi-Objective Deep Reinforcement Learning Framework

A Multi-Objective Deep Reinforcement Learning Framework

Arxiv

16+阅读 · 2018年6月27日

相关基金

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

采用pinball loss的MEE算法研究

国家自然科学基金

1+阅读 · 2013年12月31日

CTRP1促进紊流下动脉内皮屏障功能损伤及其机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

深共熔溶剂介质中镍磷合金的可控制备及电化学性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

离子液体中纤维素直接转化为5-羟甲基糠醛催化反应研究

国家自然科学基金

0+阅读 · 2012年12月31日

近似计数的算法与复杂性

国家自然科学基金

1+阅读 · 2012年12月31日

天然产物Artanomalide D及其类似物的全合成和抗肿瘤构效关系研究

国家自然科学基金

0+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

吡唑基铜配合物的合成与组装

国家自然科学基金

0+阅读 · 2011年12月31日

基于可达域计算的AWID-AWIS车辆动力学协调控制

国家自然科学基金

0+阅读 · 2010年12月31日

微信扫码咨询专知VIP会员