POMDP中基于未来依赖的基于价值的非政策评价 (Future-Dependent Value-Based Off-Policy Evaluation in POMDPs) - 专知论文

会员服务 ·

0

价值函数 · Learning · 泛函 · 估计/估计量 · Minimax ·

2022 年 7 月 26 日

Future-Dependent Value-Based Off-Policy Evaluation in POMDPs

翻译：POMDP中基于未来依赖的基于价值的非政策评价

Masatoshi Uehara,Haruka Kiyohara,Andrew Bennett,Victor Chernozhukov,Nan Jiang,Nathan Kallus,Chengchun Shi,Wen Sun

We study off-policy evaluation (OPE) for partially observable MDPs (POMDPs) with general function approximation. Existing methods such as sequential importance sampling estimators and fitted-Q evaluation suffer from the curse of horizon in POMDPs. To circumvent this problem, we develop a novel model-free OPE method by introducing future-dependent value functions that take future proxies as inputs. Future-dependent value functions play similar roles as classical value functions in fully-observable MDPs. We derive a new Bellman equation for future-dependent value functions as conditional moment equations that use history proxies as instrumental variables. We further propose a minimax learning method to learn future-dependent value functions using the new Bellman equation. We obtain the PAC result, which implies our OPE estimator is consistent as long as futures and histories contain sufficient information about latent states, and the Bellman completeness. Finally, we extend our methods to learning of dynamics and establish the connection between our approach and the well-known spectral learning methods in POMDPs.

翻译：我们研究部分可观测的 MDP (POMDP) 的离岸评估(OPE) 。现有的方法,例如顺序重要性抽样估计器和适应性-Q评价,在POMDP 中受到地平线的诅咒。为了回避这一问题,我们开发了一种新的无模式的OPE方法,引入了未来依赖性价值的功能,将未来的代理人作为投入。未来依赖性价值功能在完全可观测的 MDP 中扮演着与传统价值功能相似的作用。我们为未来依赖性价值函数制定了一个新的Bellman方程式,作为有条件的瞬间方程式,使用历史代号作为工具变量。我们进一步提出了利用新的贝尔曼方程式学习未来依赖性价值函数的微型学习方法。我们获得了PAC结果,这意味着只要未来和历史包含关于潜在状态和贝尔曼完整性的充分信息,我们的OPE 估计器就具有一致性。最后,我们将我们的方法扩大到学习动态并确定我们的方法与POMDP 中众所周知的光谱学习方法之间的联系。

0

相关内容

价值函数

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

重离子储存环CSRe上激光冷却相对论能量类锂12C3+离子束的实验研究

国家自然科学基金

0+阅读 · 2015年12月31日

纳米结构核材料离子束辐照的三维蒙特卡洛模拟

国家自然科学基金

0+阅读 · 2014年12月31日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

Bi5Fe0.5M0.5Ti3O15(M=Cr,Mn)室温多铁薄膜磁--电性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Cocycle动力学和拟周期薛定谔算子的谱

国家自然科学基金

0+阅读 · 2012年12月31日

复合壳层纳米电缆阵列的能带调控机制及其光伏特性

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

单相多铁性材料光学特性的理论研究

国家自然科学基金

0+阅读 · 2009年12月31日

d/f-电子材料的第一原理多体理论方法及其应用

国家自然科学基金

0+阅读 · 2009年12月31日

Unbiased Estimation using a Class of Diffusion Processes

Arxiv

0+阅读 · 2022年9月19日

Rewarding Episodic Visitation Discrepancy for Exploration in Reinforcement Learning

Arxiv

0+阅读 · 2022年9月19日

Towards Robust Off-Policy Evaluation via Human Inputs

Arxiv

0+阅读 · 2022年9月18日

ActiveNeRF: Learning where to See with Uncertainty Estimation

Arxiv

0+阅读 · 2022年9月18日

Data-Driven Risk-sensitive Model Predictive Control for Safe Navigation in Multi-Robot Systems

Data-Driven Risk-sensitive Model Predictive Control for Safe Navigation in Multi-Robot Systems

Arxiv

0+阅读 · 2022年9月16日

Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization

Arxiv

0+阅读 · 2022年9月15日

Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model

Arxiv

0+阅读 · 2022年9月15日

Exploiting Reward Shifting in Value-Based Deep RL

Arxiv

0+阅读 · 2022年9月15日

Robust Anytime Learning of Markov Decision Processes

Arxiv

0+阅读 · 2022年9月15日

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Arxiv

11+阅读 · 2019年11月4日

VIP会员

文章信息

相关主题

估计/估计量

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

操作系统智能体：基于多模态大模型（MLLM）的通用计算设备智能体综述

《美国太空军系统全生命周期建模、仿真与分析效能提升方案》最新84页报告

【博士论文】推进数据高效的深度学习：非参数 Transformer、主动测试与上下文学习

自主人工智能：未来战争是否将是自主化的？

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Unbiased Estimation using a Class of Diffusion Processes

Arxiv

0+阅读 · 2022年9月19日

Rewarding Episodic Visitation Discrepancy for Exploration in Reinforcement Learning

Arxiv

0+阅读 · 2022年9月19日

Towards Robust Off-Policy Evaluation via Human Inputs

Arxiv

0+阅读 · 2022年9月18日

ActiveNeRF: Learning where to See with Uncertainty Estimation

Arxiv

0+阅读 · 2022年9月18日

Data-Driven Risk-sensitive Model Predictive Control for Safe Navigation in Multi-Robot Systems

Data-Driven Risk-sensitive Model Predictive Control for Safe Navigation in Multi-Robot Systems

Arxiv

0+阅读 · 2022年9月16日

Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization

Arxiv

0+阅读 · 2022年9月15日

Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model

Arxiv

0+阅读 · 2022年9月15日

Exploiting Reward Shifting in Value-Based Deep RL

Arxiv

0+阅读 · 2022年9月15日

Robust Anytime Learning of Markov Decision Processes

Arxiv

0+阅读 · 2022年9月15日

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Arxiv

11+阅读 · 2019年11月4日

相关基金

重离子储存环CSRe上激光冷却相对论能量类锂12C3+离子束的实验研究

国家自然科学基金

0+阅读 · 2015年12月31日

纳米结构核材料离子束辐照的三维蒙特卡洛模拟

国家自然科学基金

0+阅读 · 2014年12月31日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

Bi5Fe0.5M0.5Ti3O15(M=Cr,Mn)室温多铁薄膜磁--电性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Cocycle动力学和拟周期薛定谔算子的谱

国家自然科学基金

0+阅读 · 2012年12月31日

复合壳层纳米电缆阵列的能带调控机制及其光伏特性

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

单相多铁性材料光学特性的理论研究

国家自然科学基金

0+阅读 · 2009年12月31日

d/f-电子材料的第一原理多体理论方法及其应用

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员