外部政策间隔估计 (Deeply-Debiased Off-Policy Interval Estimation) - 专知论文

会员服务 ·

0

估计/估计量 · 置信度 · HTTPS · 稳健性 · 学成 ·

2021 年 5 月 10 日

Deeply-Debiased Off-Policy Interval Estimation

翻译：外部政策间隔估计

Chengchun Shi,Runzhe Wan,Victor Chernozhukov,Rui Song

Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy. In addition to a point estimate, many applications would benefit significantly from having a confidence interval (CI) that quantifies the uncertainty of the point estimate. In this paper, we propose a novel procedure to construct an efficient, robust, and flexible CI on a target policy's value. Our method is justified by theoretical results and numerical experiments. A Python implementation of the proposed procedure is available at https://github.com/RunzheStat/D2OPE.

翻译：离政策评价以不同行为政策产生的历史数据集来学习目标政策的价值;除了点数估计外,许多应用将大大受益于一个信任间隔(CI),该间隔可以量化点数估计的不确定性;在本文中,我们提出一个新的程序,在目标政策的价值上构建一个高效、稳健和灵活的CI。我们的方法有理论结果和数字实验的正当理由。在https://github.com/RunzheStat/D2OPE上可以查阅拟议的程序的Python实施情况。

0

相关内容

估计/估计量

估计/估计量

【ICML2021】密度约束强化学习

专知会员服务

22+阅读 · 2021年6月26日

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

专知会员服务

23+阅读 · 2021年6月3日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

【AAAI2021】自校正Q学习，Self-correcting Q-Learning

专知会员服务

17+阅读 · 2020年12月4日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知会员服务

87+阅读 · 2020年8月28日

【三维物体和手部姿态估计】综述论文最新进展，Recent Advances in 3D Object and Hand Pose Estimation

【三维物体和手部姿态估计】综述论文最新进展，Recent Advances in 3D Object and Hand Pose Estimation

专知会员服务

21+阅读 · 2020年6月13日

【AISTATS2020接受论文】时空对齐，过空间和时间的最优transport（Spatio-Temporal Alignments: Optimal transport through space and time）

【AISTATS2020接受论文】时空对齐，过空间和时间的最优transport（Spatio-Temporal Alignments: Optimal transport through space and time）

专知会员服务

30+阅读 · 2020年1月11日

【Google DeepMind & 斯坦福 AAAI2020】Options of Interest Temporal Abstraction with Interest Function

【Google DeepMind & 斯坦福 AAAI2020】Options of Interest Temporal Abstraction with Interest Function

专知会员服务

5+阅读 · 2020年1月5日

【ECML-PKDD 2019】基于bagged-trees学习的可解释生存梯度提升模型（Interpretable survival gradient boosting models with bagged trees base learners）

【ECML-PKDD 2019】基于bagged-trees学习的可解释生存梯度提升模型（Interpretable survival gradient boosting models with bagged trees base learners）

专知会员服务

6+阅读 · 2019年12月1日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

【泡泡一分钟】基于运动估计的激光雷达和相机标定方法

【泡泡一分钟】基于运动估计的激光雷达和相机标定方法

泡泡机器人SLAM

25+阅读 · 2019年1月17日

meta learning 17年：MAML SNAIL

meta learning 17年：MAML SNAIL

CreateAMind

11+阅读 · 2019年1月2日

Disentangled的假设的探讨

Disentangled的假设的探讨

CreateAMind

9+阅读 · 2018年12月10日

Hierarchical Disentangled Representations

Hierarchical Disentangled Representations

CreateAMind

4+阅读 · 2018年4月15日

lightgbm algorithm case of kaggle（上）

lightgbm algorithm case of kaggle（上）

R语言中文社区

8+阅读 · 2018年3月20日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

强化学习 cartpole_a3c

强化学习 cartpole_a3c

CreateAMind

9+阅读 · 2017年7月21日

已删除

将门创投

4+阅读 · 2017年7月7日

Learning from an Exploring Demonstrator: Optimal Reward Estimation for Bandits

Arxiv

0+阅读 · 2021年6月28日

Whittle estimation with (quasi-)analytic wavelets

Arxiv

0+阅读 · 2021年6月28日

Query Complexity of Least Absolute Deviation Regression via Robust Uniform Convergence

Arxiv

0+阅读 · 2021年6月28日

VID-Fusion: Robust Visual-Inertial-Dynamics Odometry for Accurate External Force Estimation

Arxiv

0+阅读 · 2021年6月28日

Nonparametric estimation of continuous DPPs with kernel methods

Arxiv

0+阅读 · 2021年6月27日

On High Dimensional Covariate Adjustment for Estimating Causal Effects in Randomized Trials with Survival Outcomes

Arxiv

0+阅读 · 2021年6月25日

Bootstrap confidence intervals for multiple change points based on moving sum procedures

Arxiv

0+阅读 · 2021年6月24日

Faster Policy Learning with Continuous-Time Gradients

Arxiv

0+阅读 · 2021年6月24日

R-LINS: A Robocentric Lidar-Inertial State Estimator for Robust and Efficient Navigation

R-LINS: A Robocentric Lidar-Inertial State Estimator for Robust and Efficient Navigation

Arxiv

3+阅读 · 2019年8月22日

Parameter Space Noise for Exploration

Arxiv

3+阅读 · 2018年1月31日

VIP会员

文章信息

相关主题

估计/估计量

相关VIP内容

【ICML2021】密度约束强化学习

专知会员服务

22+阅读 · 2021年6月26日

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

专知会员服务

23+阅读 · 2021年6月3日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

【AAAI2021】自校正Q学习，Self-correcting Q-Learning

专知会员服务

17+阅读 · 2020年12月4日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知会员服务

87+阅读 · 2020年8月28日

【三维物体和手部姿态估计】综述论文最新进展，Recent Advances in 3D Object and Hand Pose Estimation

【三维物体和手部姿态估计】综述论文最新进展，Recent Advances in 3D Object and Hand Pose Estimation

专知会员服务

21+阅读 · 2020年6月13日

【AISTATS2020接受论文】时空对齐，过空间和时间的最优transport（Spatio-Temporal Alignments: Optimal transport through space and time）

【AISTATS2020接受论文】时空对齐，过空间和时间的最优transport（Spatio-Temporal Alignments: Optimal transport through space and time）

专知会员服务

30+阅读 · 2020年1月11日

【Google DeepMind & 斯坦福 AAAI2020】Options of Interest Temporal Abstraction with Interest Function

【Google DeepMind & 斯坦福 AAAI2020】Options of Interest Temporal Abstraction with Interest Function

专知会员服务

5+阅读 · 2020年1月5日

【ECML-PKDD 2019】基于bagged-trees学习的可解释生存梯度提升模型（Interpretable survival gradient boosting models with bagged trees base learners）

【ECML-PKDD 2019】基于bagged-trees学习的可解释生存梯度提升模型（Interpretable survival gradient boosting models with bagged trees base learners）

专知会员服务

6+阅读 · 2019年12月1日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

热门VIP内容

开通专知VIP会员享更多权益服务

《物联网（IoT）中的无人机通信高效控制》135页

《在GNSS信号降级环境中利用共识实现无人机集群稳健协调》

中程单向攻击无人机的战略意义：俄乌战争启示

《面向无人机集群的避障动态传感器覆盖算法》最新38页

相关资讯

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

【泡泡一分钟】基于运动估计的激光雷达和相机标定方法

【泡泡一分钟】基于运动估计的激光雷达和相机标定方法

泡泡机器人SLAM

25+阅读 · 2019年1月17日

meta learning 17年：MAML SNAIL

meta learning 17年：MAML SNAIL

CreateAMind

11+阅读 · 2019年1月2日

Disentangled的假设的探讨

Disentangled的假设的探讨

CreateAMind

9+阅读 · 2018年12月10日

Hierarchical Disentangled Representations

Hierarchical Disentangled Representations

CreateAMind

4+阅读 · 2018年4月15日

lightgbm algorithm case of kaggle（上）

lightgbm algorithm case of kaggle（上）

R语言中文社区

8+阅读 · 2018年3月20日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

强化学习 cartpole_a3c

强化学习 cartpole_a3c

CreateAMind

9+阅读 · 2017年7月21日

已删除

将门创投

4+阅读 · 2017年7月7日

相关论文

Learning from an Exploring Demonstrator: Optimal Reward Estimation for Bandits

Arxiv

0+阅读 · 2021年6月28日

Whittle estimation with (quasi-)analytic wavelets

Arxiv

0+阅读 · 2021年6月28日

Query Complexity of Least Absolute Deviation Regression via Robust Uniform Convergence

Arxiv

0+阅读 · 2021年6月28日

VID-Fusion: Robust Visual-Inertial-Dynamics Odometry for Accurate External Force Estimation

Arxiv

0+阅读 · 2021年6月28日

Nonparametric estimation of continuous DPPs with kernel methods

Arxiv

0+阅读 · 2021年6月27日

On High Dimensional Covariate Adjustment for Estimating Causal Effects in Randomized Trials with Survival Outcomes

Arxiv

0+阅读 · 2021年6月25日

Bootstrap confidence intervals for multiple change points based on moving sum procedures

Arxiv

0+阅读 · 2021年6月24日

Faster Policy Learning with Continuous-Time Gradients

Arxiv

0+阅读 · 2021年6月24日

R-LINS: A Robocentric Lidar-Inertial State Estimator for Robust and Efficient Navigation

R-LINS: A Robocentric Lidar-Inertial State Estimator for Robust and Efficient Navigation

Arxiv

3+阅读 · 2019年8月22日

Parameter Space Noise for Exploration

Arxiv

3+阅读 · 2018年1月31日

微信扫码咨询专知VIP会员