改进蒙特卡罗评估方法：离线数据的应用 (Improving Monte Carlo Evaluation with Offline Data) - 专知论文

会员服务 ·

0

估计/估计量 · 蒙特卡罗 · 在线 · 可约的 · 样本 ·

2023 年 3 月 23 日

Improving Monte Carlo Evaluation with Offline Data

翻译：改进蒙特卡罗评估方法：离线数据的应用

Shuze Liu,Shangtong Zhang

Monte Carlo (MC) methods are the most widely used methods to estimate the performance of a policy. Given an interested policy, MC methods give estimates by repeatedly running this policy to collect samples and taking the average of the outcomes. Samples collected during this process are called online samples. To get an accurate estimate, MC methods consume massive online samples. When online samples are expensive, e.g., online recommendations and inventory management, we want to reduce the number of online samples while achieving the same estimate accuracy. To this end, we use off-policy MC methods that evaluate the interested policy by running a different policy called behavior policy. We design a tailored behavior policy such that the variance of the off-policy MC estimator is provably smaller than the ordinary MC estimator. Importantly, this tailored behavior policy can be efficiently learned from existing offline data, i,e., previously logged data, which are much cheaper than online samples. With reduced variance, our off-policy MC method requires fewer online samples to evaluate the performance of a policy compared with the ordinary MC method. Moreover, our off-policy MC estimator is always unbiased.

翻译：蒙特卡罗(Monte Carlo, MC)方法是最常用的评估政策性能的方法。对于给定的政策，MC方法通过反复运行该政策来收集样本，并取得结果的平均值来给出估计。在此过程中收集到的样本称为在线样本。为了得到准确的估计，MC方法需要消耗大量的在线样本。当在线样本比较昂贵时，例如在线推荐和库存管理，我们希望在保持相同的估计准确性的同时减少在线样本数量。为此，我们使用离线MC方法通过运行不同的策略(行为策略)对所需策略进行评估。我们设计出一种特定的行为策略，使得离线MC估计器的方差比普通MC估计器的方差小。重要的是，这种定制的行为策略可以从现有的离线数据(即先前记录的数据)中高效地学习，而这比在线样本要便宜得多。通过减小方差，我们的离线MC方法比普通MC方法需要更少的在线样本来评估政策的表现。此外，我们的离线MC估计器始终是无偏的。

0

相关内容

估计/估计量

估计/估计量

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

专知会员服务

63+阅读 · 2023年2月15日

【ICDM 2022教程】图挖掘中的公平性:度量、算法和应用

【ICDM 2022教程】图挖掘中的公平性:度量、算法和应用

专知会员服务

28+阅读 · 2022年12月26日

【CIKM2021】基于检索的个性化聊天机器人模型IMPChat

专知会员服务

17+阅读 · 2021年8月25日

【剑桥大学】统计因果关系的决策理论基础，Decision-theoretic foundations for statistical causality

【剑桥大学】统计因果关系的决策理论基础，Decision-theoretic foundations for statistical causality

专知会员服务

48+阅读 · 2020年5月5日

【ACL2020】对抗性文本生成，Improving Adversarial Text Generation

专知会员服务

52+阅读 · 2020年5月5日

【Google-Mila】你的GAN实际上是一个基于能量的模型，你应该使用鉴别器驱动的潜在采样，Your GAN is Secretly an Energy-based Model and You Should Use Discriminator Driven Latent Sampling

【Google-Mila】你的GAN实际上是一个基于能量的模型，你应该使用鉴别器驱动的潜在采样，Your GAN is Secretly an Energy-based Model and You Should Use Discriminator Driven Latent Sampling

专知会员服务

30+阅读 · 2020年3月28日

最大均方差正则化贝叶斯神经网络，Bayesian Neural Networks With Maximum Mean Discrepancy Regularization

最大均方差正则化贝叶斯神经网络，Bayesian Neural Networks With Maximum Mean Discrepancy Regularization

专知会员服务

54+阅读 · 2020年3月5日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文推荐】最新八篇推荐系统相关论文—亿级商品嵌入、主动学习、树深度模型、知识图谱、注意力感知、矩阵分解、神经个性化嵌入

【论文推荐】最新八篇推荐系统相关论文—亿级商品嵌入、主动学习、树深度模型、知识图谱、注意力感知、矩阵分解、神经个性化嵌入

专知

15+阅读 · 2018年6月15日

强化学习初探 - 从多臂老虎机问题说起

强化学习初探 - 从多臂老虎机问题说起

专知

10+阅读 · 2018年4月3日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

间接优化的高效Monte Carlo声传播研究

国家自然科学基金

0+阅读 · 2017年12月31日

函数数据变换模型及降维方法的研究

国家自然科学基金

1+阅读 · 2015年12月31日

复杂数据下含指标项半参数模型结构的统计推断及应用

国家自然科学基金

0+阅读 · 2014年12月31日

基于多粒度粗糙集的风险投资项目选择决策分析研究

国家自然科学基金

2+阅读 · 2013年12月31日

半参数有效性工具变量理论：选择、估计、检验及其应用

国家自然科学基金

0+阅读 · 2013年12月31日

增量协同过滤模型研究

国家自然科学基金

0+阅读 · 2012年12月31日

随机模糊环境下的数据包络分析

国家自然科学基金

0+阅读 · 2012年12月31日

区间删失数据的半参数回归模型的有效估计方法

国家自然科学基金

0+阅读 · 2012年12月31日

基于贝叶斯推理的模糊逻辑强化学习模型研究

国家自然科学基金

18+阅读 · 2012年12月31日

空间经济计量模型检验中Bootstrap方法有效性研究

国家自然科学基金

0+阅读 · 2008年12月31日

Improving ChatGPT Prompt for Code Generation

Arxiv

0+阅读 · 2023年5月15日

Evaluating Self-Supervised Learning via Risk Decomposition

Arxiv

0+阅读 · 2023年5月13日

Smoothed empirical likelihood estimation and automatic variable selection for an expectile high-dimensional model with possibly missing response variable

Smoothed empirical likelihood estimation and automatic variable selection for an expectile high-dimensional model with possibly missing response variable

Arxiv

0+阅读 · 2023年5月12日

Differentially Private Set-Based Estimation Using Zonotopes

Arxiv

0+阅读 · 2023年5月12日

Sequential model correction for nonlinear inverse problems

Arxiv

0+阅读 · 2023年5月12日

Adaptive Privacy-Preserving Coded Computing With Hierarchical Task Partitioning

Arxiv

0+阅读 · 2023年5月11日

Correlation visualization under missing values: a comparison between imputation and direct parameter estimation methods

Arxiv

0+阅读 · 2023年5月10日

Privacy-Preserving Recommender Systems with Synthetic Query Generation using Differentially Private Large Language Models

Arxiv

0+阅读 · 2023年5月10日

Fast Distributed Inference Serving for Large Language Models

Arxiv

0+阅读 · 2023年5月10日

Causal Embeddings for Recommendation

Arxiv

23+阅读 · 2018年8月3日

VIP会员

文章信息

相关主题

估计/估计量

相关VIP内容

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

专知会员服务

63+阅读 · 2023年2月15日

【ICDM 2022教程】图挖掘中的公平性:度量、算法和应用

【ICDM 2022教程】图挖掘中的公平性:度量、算法和应用

专知会员服务

28+阅读 · 2022年12月26日

【CIKM2021】基于检索的个性化聊天机器人模型IMPChat

专知会员服务

17+阅读 · 2021年8月25日

【剑桥大学】统计因果关系的决策理论基础，Decision-theoretic foundations for statistical causality

【剑桥大学】统计因果关系的决策理论基础，Decision-theoretic foundations for statistical causality

专知会员服务

48+阅读 · 2020年5月5日

【ACL2020】对抗性文本生成，Improving Adversarial Text Generation

专知会员服务

52+阅读 · 2020年5月5日

【Google-Mila】你的GAN实际上是一个基于能量的模型，你应该使用鉴别器驱动的潜在采样，Your GAN is Secretly an Energy-based Model and You Should Use Discriminator Driven Latent Sampling

【Google-Mila】你的GAN实际上是一个基于能量的模型，你应该使用鉴别器驱动的潜在采样，Your GAN is Secretly an Energy-based Model and You Should Use Discriminator Driven Latent Sampling

专知会员服务

30+阅读 · 2020年3月28日

最大均方差正则化贝叶斯神经网络，Bayesian Neural Networks With Maximum Mean Discrepancy Regularization

最大均方差正则化贝叶斯神经网络，Bayesian Neural Networks With Maximum Mean Discrepancy Regularization

专知会员服务

54+阅读 · 2020年3月5日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

新书册《几何深度学习的数学基础》

中程单向攻击无人机的战略意义：俄乌战争启示

在无标注条件下适配视觉—语言模型：全面综述

面向视觉语言模型的持续学习：遗忘之外的综述与分类体系

相关资讯

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文推荐】最新八篇推荐系统相关论文—亿级商品嵌入、主动学习、树深度模型、知识图谱、注意力感知、矩阵分解、神经个性化嵌入

【论文推荐】最新八篇推荐系统相关论文—亿级商品嵌入、主动学习、树深度模型、知识图谱、注意力感知、矩阵分解、神经个性化嵌入

专知

15+阅读 · 2018年6月15日

强化学习初探 - 从多臂老虎机问题说起

强化学习初探 - 从多臂老虎机问题说起

专知

10+阅读 · 2018年4月3日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Improving ChatGPT Prompt for Code Generation

Arxiv

0+阅读 · 2023年5月15日

Evaluating Self-Supervised Learning via Risk Decomposition

Arxiv

0+阅读 · 2023年5月13日

Smoothed empirical likelihood estimation and automatic variable selection for an expectile high-dimensional model with possibly missing response variable

Smoothed empirical likelihood estimation and automatic variable selection for an expectile high-dimensional model with possibly missing response variable

Arxiv

0+阅读 · 2023年5月12日

Differentially Private Set-Based Estimation Using Zonotopes

Arxiv

0+阅读 · 2023年5月12日

Sequential model correction for nonlinear inverse problems

Arxiv

0+阅读 · 2023年5月12日

Adaptive Privacy-Preserving Coded Computing With Hierarchical Task Partitioning

Arxiv

0+阅读 · 2023年5月11日

Correlation visualization under missing values: a comparison between imputation and direct parameter estimation methods

Arxiv

0+阅读 · 2023年5月10日

Privacy-Preserving Recommender Systems with Synthetic Query Generation using Differentially Private Large Language Models

Arxiv

0+阅读 · 2023年5月10日

Fast Distributed Inference Serving for Large Language Models

Arxiv

0+阅读 · 2023年5月10日

Causal Embeddings for Recommendation

Arxiv

23+阅读 · 2018年8月3日

相关基金

间接优化的高效Monte Carlo声传播研究

国家自然科学基金

0+阅读 · 2017年12月31日

函数数据变换模型及降维方法的研究

国家自然科学基金

1+阅读 · 2015年12月31日

复杂数据下含指标项半参数模型结构的统计推断及应用

国家自然科学基金

0+阅读 · 2014年12月31日

基于多粒度粗糙集的风险投资项目选择决策分析研究

国家自然科学基金

2+阅读 · 2013年12月31日

半参数有效性工具变量理论：选择、估计、检验及其应用

国家自然科学基金

0+阅读 · 2013年12月31日

增量协同过滤模型研究

国家自然科学基金

0+阅读 · 2012年12月31日

随机模糊环境下的数据包络分析

国家自然科学基金

0+阅读 · 2012年12月31日

区间删失数据的半参数回归模型的有效估计方法

国家自然科学基金

0+阅读 · 2012年12月31日

基于贝叶斯推理的模糊逻辑强化学习模型研究

国家自然科学基金

18+阅读 · 2012年12月31日

空间经济计量模型检验中Bootstrap方法有效性研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员