最优化的内幕强盗和通过回归神器可实现的Knapsacks (Optimal Contextual Bandits with Knapsacks under Realizibility via Regression Oracles) - 专知论文

会员服务 ·

0

上下文赌博机/上下文老虎机 · 赌博机/老虎机 · 泛函 · 广义函数 · 优化器 ·

2022 年 10 月 21 日

Optimal Contextual Bandits with Knapsacks under Realizibility via Regression Oracles

翻译：最优化的内幕强盗和通过回归神器可实现的Knapsacks

Yuxuan Han,Jialin Zeng,Yang Wang,Yang Xiang,Jiheng Zhang

We study the stochastic contextual bandit with knapsacks (CBwK) problem, where each action, taken upon a context, not only leads to a random reward but also costs a random resource consumption in a vector form. The challenge is to maximize the total reward without violating the budget for each resource. We study this problem under a general realizability setting where the expected reward and expected cost are functions of contexts and actions in some given general function classes $\mathcal{F}$ and $\mathcal{G}$, respectively. Existing works on CBwK are restricted to the linear function class since they use UCB-type algorithms, which heavily rely on the linear form and thus are difficult to extend to general function classes. Motivated by online regression oracles that have been successfully applied to contextual bandits, we propose the first universal and optimal algorithmic framework for CBwK by reducing it to online regression. We also establish the lower regret bound to show the optimality of our algorithm for a variety of function classes.

翻译：我们用 knapsacks (CBwK) 来研究背景上的土匪问题, 每一个行动都是在某种背景下采取的, 不仅导致随机的奖励, 而且还以矢量形式花费随机的资源消耗。挑战是如何在不侵犯每种资源的预算的情况下最大限度地获得全部的奖励。我们在一个总体的可变性环境下研究这一问题, 因为在一般功能类别中, 所预期的奖励和预期成本分别是环境和行动功能的函数 $\ mathcal{F} $ 和$\ mathcal{G} $。 CBwK 的现有工程仅限于线性功能类别, 因为它们使用非常依赖线性形式的UCB型算法, 因而难以扩展到普通功能类别。我们受已成功应用到环境强盗的在线回归或触法驱动, 我们提出CBWK 的第一个普遍和最佳的算法框架, 将其降低到在线回归。我们还设定了较低的遗憾约束, 以显示各种功能类的算法的最佳性。

0

相关内容

上下文赌博机/上下文老虎机

上下文赌博机/上下文老虎机

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

重金属离子胁迫下花斑裸鲤钙调蛋白磷酸酶(Calcineurin)的应答及其分子调节机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

蛋白激酶调控拟南芥响应低温胁迫的分子机理

国家自然科学基金

0+阅读 · 2013年12月31日

Act1与IL-17/IL-17R串话在口腔扁平苔藓发生发展中作用的研究

国家自然科学基金

0+阅读 · 2013年12月31日

Prohibitin1在胆管癌中的作用及分子机制

国家自然科学基金

0+阅读 · 2013年12月31日

NLRP3炎性小体介导同型半胱氨酸诱导动脉粥样硬化炎症反应的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

HMGB1在白血病细胞凋亡和自噬调控机制的研究

国家自然科学基金

0+阅读 · 2012年12月31日

MCM3-SYF2复合物对cyclin D1-CDKs调节在星形胶质细胞炎症激活中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

去酰基化ghrelin改善脂肪组织炎症所致胰岛素抵抗的机制- - 调节性T细胞的作用

国家自然科学基金

0+阅读 · 2011年12月31日

遍历哈密顿系统的谱理论

国家自然科学基金

0+阅读 · 2009年12月31日

垂直磁记录介质中飞秒激光诱导的磁性软化研究

国家自然科学基金

0+阅读 · 2008年12月31日

An Improved Algorithm For Online Reranking

Arxiv

0+阅读 · 2022年12月4日

Model Selection in Contextual Stochastic Bandit Problems

Arxiv

0+阅读 · 2022年12月4日

Approximate Factor Models with Weaker Loadings

Arxiv

0+阅读 · 2022年12月4日

A Unified Quantum Algorithm Framework for Estimating Properties of Discrete Probability Distributions

Arxiv

0+阅读 · 2022年12月3日

Pandora's Problem with Nonobligatory Inspection: Optimal Structure and a PTAS

Arxiv

0+阅读 · 2022年12月3日

Pandora Box Problem with Nonobligatory Inspection: Hardness and Approximation Scheme

Arxiv

0+阅读 · 2022年12月3日

Testing Linear Operator Constraints in Functional Response Regression with Incomplete Response Functions

Arxiv

0+阅读 · 2022年12月2日

Decision Market Based Learning For Multi-agent Contextual Bandit Problems

Arxiv

0+阅读 · 2022年12月1日

Augmenting Basis Sets by Normalizing Flows

Arxiv

0+阅读 · 2022年11月30日

Combined numerical methods for solving time-varying semilinear differential-algebraic equations with the use of spectral projectors and recalculation

Arxiv

0+阅读 · 2022年11月27日

VIP会员

文章信息

相关主题

上下文赌博机/上下文老虎机

赌博机/老虎机

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《乌克兰无人机产业：志愿者与政策在构建新兴无人机产业中的协同作用》最新报告

《人工智能辅助决策中的数据可视化：系统性综述》

人工智能驱动弹药制造现代化：美国陆军转型之路

《敏捷作战部署中枢纽-辐条基地选址优化研究》80页

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

相关论文

An Improved Algorithm For Online Reranking

Arxiv

0+阅读 · 2022年12月4日

Model Selection in Contextual Stochastic Bandit Problems

Arxiv

0+阅读 · 2022年12月4日

Approximate Factor Models with Weaker Loadings

Arxiv

0+阅读 · 2022年12月4日

A Unified Quantum Algorithm Framework for Estimating Properties of Discrete Probability Distributions

Arxiv

0+阅读 · 2022年12月3日

Pandora's Problem with Nonobligatory Inspection: Optimal Structure and a PTAS

Arxiv

0+阅读 · 2022年12月3日

Pandora Box Problem with Nonobligatory Inspection: Hardness and Approximation Scheme

Arxiv

0+阅读 · 2022年12月3日

Testing Linear Operator Constraints in Functional Response Regression with Incomplete Response Functions

Arxiv

0+阅读 · 2022年12月2日

Decision Market Based Learning For Multi-agent Contextual Bandit Problems

Arxiv

0+阅读 · 2022年12月1日

Augmenting Basis Sets by Normalizing Flows

Arxiv

0+阅读 · 2022年11月30日

Combined numerical methods for solving time-varying semilinear differential-algebraic equations with the use of spectral projectors and recalculation

Arxiv

0+阅读 · 2022年11月27日

相关基金

重金属离子胁迫下花斑裸鲤钙调蛋白磷酸酶(Calcineurin)的应答及其分子调节机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

蛋白激酶调控拟南芥响应低温胁迫的分子机理

国家自然科学基金

0+阅读 · 2013年12月31日

Act1与IL-17/IL-17R串话在口腔扁平苔藓发生发展中作用的研究

国家自然科学基金

0+阅读 · 2013年12月31日

Prohibitin1在胆管癌中的作用及分子机制

国家自然科学基金

0+阅读 · 2013年12月31日

NLRP3炎性小体介导同型半胱氨酸诱导动脉粥样硬化炎症反应的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

HMGB1在白血病细胞凋亡和自噬调控机制的研究

国家自然科学基金

0+阅读 · 2012年12月31日

MCM3-SYF2复合物对cyclin D1-CDKs调节在星形胶质细胞炎症激活中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

去酰基化ghrelin改善脂肪组织炎症所致胰岛素抵抗的机制- - 调节性T细胞的作用

国家自然科学基金

0+阅读 · 2011年12月31日

遍历哈密顿系统的谱理论

国家自然科学基金

0+阅读 · 2009年12月31日

垂直磁记录介质中飞秒激光诱导的磁性软化研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员