在有常规函数接近的 Markov 游戏中进行离线学习 (Offline Learning in Markov Games with General Function Approximation) - 专知论文

会员服务 ·

0

Markov · 广义函数 · 近似 · Learning · 泛函 ·

2023 年 2 月 6 日

Offline Learning in Markov Games with General Function Approximation

翻译：在有常规函数接近的 Markov 游戏中进行离线学习

Yuheng Zhang,Yu Bai,Nan Jiang

We study offline multi-agent reinforcement learning (RL) in Markov games, where the goal is to learn an approximate equilibrium -- such as Nash equilibrium and (Coarse) Correlated Equilibrium -- from an offline dataset pre-collected from the game. Existing works consider relatively restricted tabular or linear models and handle each equilibria separately. In this work, we provide the first framework for sample-efficient offline learning in Markov games under general function approximation, handling all 3 equilibria in a unified manner. By using Bellman-consistent pessimism, we obtain interval estimation for policies' returns, and use both the upper and the lower bounds to obtain a relaxation on the gap of a candidate policy, which becomes our optimization objective. Our results generalize prior works and provide several additional insights. Importantly, we require a data coverage condition that improves over the recently proposed "unilateral concentrability". Our condition allows selective coverage of deviation policies that optimally trade-off between their greediness (as approximate best responses) and coverage, and we show scenarios where this leads to significantly better guarantees. As a new connection, we also show how our algorithmic framework can subsume seemingly different solution concepts designed for the special case of two-player zero-sum games.

翻译：我们在Markov游戏中研究离线多试剂强化学习(RL),目的是从游戏前收集的离线数据集中学习近似平衡 -- -- 例如纳什平衡和(粗)Cor-Col-alumlium -- -- 从游戏前收集的离线数据集中学习近似平衡 -- -- 例如纳什平衡和(粗)Cor-Cor-alumlium。现有的作品考虑相对有限的表格式或线性模型,并分别处理每种平衡。在这项工作中,我们为在通用功能近似下,在Markov游戏中抽样高效的离线学习提供了第一个框架,以统一的方式处理所有3种平衡。通过使用Bellman-一致的悲观主义,我们获得了政策回报的间隔估计,并利用上下限和下限,以获得对候选政策差距的宽放,这已成为我们的优化目标。我们的成果一般化了先前的工作,并提供了另外的见解。重要的是,我们需要一个数据覆盖条件,比最近提出的“单方相可调”的组合。我们的条件允许有选择地覆盖偏差政策的范围,在它们的贪婪(作为最好的回应)和覆盖范围,我们所设计的分数的分解法框架。

0

相关内容

Markov

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日

深度强化学习策略梯度教程，53页ppt

深度强化学习策略梯度教程，53页ppt

专知会员服务

184+阅读 · 2020年2月1日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

量化金融强化学习论文集合

量化金融强化学习论文集合

专知

14+阅读 · 2019年12月18日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

F-actin结合蛋白在维甲酸诱导的舌肌发育不良中的作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

Cell-in-cell介导非易感细胞病毒感染及其免疫逃逸机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Schrodinger-Poisson方程的若干问题研究

国家自然科学基金

1+阅读 · 2012年12月31日

Markov决策过程值函数逼近的基函数自动构造

国家自然科学基金

1+阅读 · 2012年12月31日

巨磁致伸缩材料中磁机械效应和磁致伸缩

国家自然科学基金

0+阅读 · 2012年12月31日

SrxBa1-xNb2O6纳米陶瓷与薄膜的电卡效应

国家自然科学基金

0+阅读 · 2012年12月31日

免疫佐剂激发PCV-2感染的分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

反铁电体中的热释电效应和电卡效应研究

国家自然科学基金

0+阅读 · 2008年12月31日

Learning Flow Functions from Data with Applications to Nonlinear Oscillators

Arxiv

0+阅读 · 2023年3月29日

Geometric Convergence of Distributed Heavy-Ball Nash Equilibrium Algorithm over Time-Varying Digraphs with Unconstrained Actions

Arxiv

0+阅读 · 2023年3月29日

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Arxiv

0+阅读 · 2023年3月28日

A General Unfolding Speech Enhancement Method Motivated by Taylor's Theorem

Arxiv

0+阅读 · 2023年3月28日

Online Learning for Equilibrium Pricing in Markets under Incomplete Information

Arxiv

0+阅读 · 2023年3月28日

Efficient Activation Function Optimization through Surrogate Modeling

Arxiv

0+阅读 · 2023年3月27日

Sequential Knockoffs for Variable Selection in Reinforcement Learning

Arxiv

0+阅读 · 2023年3月24日

Higher order time discretization method for a class of semilinear stochastic partial differential equations with multiplicative noise

Arxiv

0+阅读 · 2023年3月24日

Safety-Critical Coordination for Cooperative Legged Locomotion via Control Barrier Functions

Arxiv

0+阅读 · 2023年3月23日

Multi-Agent Cooperative Bidding Games for Multi-Objective Optimization in e-Commercial Sponsored Search

Arxiv

12+阅读 · 2021年6月8日

VIP会员

文章信息

相关主题

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日

深度强化学习策略梯度教程，53页ppt

深度强化学习策略梯度教程，53页ppt

专知会员服务

184+阅读 · 2020年2月1日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《战区安全决策课程体系》最新244页

《"无人机航母"原型平台》

任务规划与地形分析：现代复杂环境作战导航体系

《攻击场景描述形式化模型研究》

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

量化金融强化学习论文集合

量化金融强化学习论文集合

专知

14+阅读 · 2019年12月18日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Learning Flow Functions from Data with Applications to Nonlinear Oscillators

Arxiv

0+阅读 · 2023年3月29日

Geometric Convergence of Distributed Heavy-Ball Nash Equilibrium Algorithm over Time-Varying Digraphs with Unconstrained Actions

Arxiv

0+阅读 · 2023年3月29日

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Arxiv

0+阅读 · 2023年3月28日

A General Unfolding Speech Enhancement Method Motivated by Taylor's Theorem

Arxiv

0+阅读 · 2023年3月28日

Online Learning for Equilibrium Pricing in Markets under Incomplete Information

Arxiv

0+阅读 · 2023年3月28日

Efficient Activation Function Optimization through Surrogate Modeling

Arxiv

0+阅读 · 2023年3月27日

Sequential Knockoffs for Variable Selection in Reinforcement Learning

Arxiv

0+阅读 · 2023年3月24日

Higher order time discretization method for a class of semilinear stochastic partial differential equations with multiplicative noise

Arxiv

0+阅读 · 2023年3月24日

Safety-Critical Coordination for Cooperative Legged Locomotion via Control Barrier Functions

Arxiv

0+阅读 · 2023年3月23日

Multi-Agent Cooperative Bidding Games for Multi-Objective Optimization in e-Commercial Sponsored Search

Arxiv

12+阅读 · 2021年6月8日

相关基金

F-actin结合蛋白在维甲酸诱导的舌肌发育不良中的作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

Cell-in-cell介导非易感细胞病毒感染及其免疫逃逸机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Schrodinger-Poisson方程的若干问题研究

国家自然科学基金

1+阅读 · 2012年12月31日

Markov决策过程值函数逼近的基函数自动构造

国家自然科学基金

1+阅读 · 2012年12月31日

巨磁致伸缩材料中磁机械效应和磁致伸缩

国家自然科学基金

0+阅读 · 2012年12月31日

SrxBa1-xNb2O6纳米陶瓷与薄膜的电卡效应

国家自然科学基金

0+阅读 · 2012年12月31日

免疫佐剂激发PCV-2感染的分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

反铁电体中的热释电效应和电卡效应研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员