全部挤压: 新动动画和线性内地土匪自热带 (Squeeze All: Novel Estimator and Self-Normalized Bound for Linear Contextual Bandits) - 专知论文

会员服务 ·

0

上下文赌博机/上下文老虎机 · 估计/估计量 · 赌博机/老虎机 · 线性的 · 情景 ·

2022 年 6 月 16 日

Squeeze All: Novel Estimator and Self-Normalized Bound for Linear Contextual Bandits

翻译：全部挤压: 新动动画和线性内地土匪自热带

Wonyoung Kim,Min-hwan Oh,Myunghee Cho Paik

from arxiv, 27 pages including Appendix

We propose a novel algorithm for linear contextual bandits with $O(\sqrt{dT \log T})$ regret bound, where $d$ is the dimension of contexts and $T$ is the time horizon. Our proposed algorithm is equipped with a novel estimator in which exploration is embedded through explicit randomization. Depending on the randomization, our proposed estimator takes contribution either from contexts of all arms or from selected contexts. We establish a self-normalized bound for our estimator, which allows a novel decomposition of the cumulative regret into additive dimension-dependent terms instead of multiplicative terms. We also prove a novel lower bound of $\Omega(\sqrt{dT})$ under our problem setting. Hence, the regret of our proposed algorithm matches the lower bound up to logarithmic factors. The numerical experiments support the theoretical guarantees and show that our proposed method outperforms the existing linear bandit algorithms.

翻译：我们为线性背景强盗建议了一个新奇的算法,使用$O(\\ sqrt{dTT\log T}) 表示歉意, 美元是环境的维度, 美元是时间范围。我们提议的算法配备了一个新颖的测算器, 通过明确的随机化嵌入勘探。取决于随机化, 我们提议的测算器会从所有手臂或选定环境的背景中做出贡献。我们为我们的测算仪建立一个自我标准化的定线, 允许将累积的差分解成基于添加维度的术语, 而不是乘数性术语。在我们的问题设置下, 我们还证明了一个小于$\ Omega( sqrt{dT}) 的新颖的下限。因此, 我们提议的算法的遗憾匹配了较低约束到对线性系数。数字实验支持理论保证, 并显示我们提议的算法比现有的线式带算法要优。

0

相关内容

上下文赌博机/上下文老虎机

上下文赌博机/上下文老虎机

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

Anderson型多酸的不对称修饰及可控组装研究

国家自然科学基金

1+阅读 · 2014年12月31日

新泛素化修饰因子对Hedgehog信号通路调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Long non-coding RNA MEG3分子对胶质瘤干细胞调控作用的研究

国家自然科学基金

0+阅读 · 2013年12月31日

IGF-1R糖基化修饰通过调控内质网应激影响卵巢癌细胞生物学行为的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

Centroids Matching: an efficient Continual Learning approach operating in the embedding space

Arxiv

0+阅读 · 2022年8月3日

S-estimation in Linear Models with Structured Covariance Matrices

Arxiv

0+阅读 · 2022年8月3日

Debiasing In-Sample Policy Performance for Small-Data, Large-Scale Optimization

Arxiv

0+阅读 · 2022年8月2日

An unbiased minimum variance non-parametric analytic and likelihood estimator for discrete and continuous score spaces

Arxiv

0+阅读 · 2022年8月2日

Numerical identification of initial temperatures in heat equation with dynamic boundary conditions

Arxiv

0+阅读 · 2022年8月1日

VIP会员

文章信息

相关主题

上下文赌博机/上下文老虎机

估计/估计量

赌博机/老虎机

相关VIP内容

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《基于AI的动态任务分配策略实现多智能体系统有意义人类控制》报告

《超越连接：AI驱动网络未来愿景》最新报告

人工智能赋能多域作战：能力与挑战

《战场空间决策优势：AI基础与应用研究》总结报告

相关资讯

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

相关论文

Centroids Matching: an efficient Continual Learning approach operating in the embedding space

Arxiv

0+阅读 · 2022年8月3日

S-estimation in Linear Models with Structured Covariance Matrices

Arxiv

0+阅读 · 2022年8月3日

Debiasing In-Sample Policy Performance for Small-Data, Large-Scale Optimization

Arxiv

0+阅读 · 2022年8月2日

An unbiased minimum variance non-parametric analytic and likelihood estimator for discrete and continuous score spaces

Arxiv

0+阅读 · 2022年8月2日

Numerical identification of initial temperatures in heat equation with dynamic boundary conditions

Arxiv

0+阅读 · 2022年8月1日

相关基金

Anderson型多酸的不对称修饰及可控组装研究

国家自然科学基金

1+阅读 · 2014年12月31日

新泛素化修饰因子对Hedgehog信号通路调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Long non-coding RNA MEG3分子对胶质瘤干细胞调控作用的研究

国家自然科学基金

0+阅读 · 2013年12月31日

IGF-1R糖基化修饰通过调控内质网应激影响卵巢癌细胞生物学行为的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员