沙沙起伏强盗 (Stochastic Rising Bandits) - 专知论文

会员服务 ·

0

赌博机/老虎机 · CASE · state-of-the-art · REST · 在线 ·

2022 年 12 月 7 日

Stochastic Rising Bandits

翻译：沙沙起伏强盗

Alberto Maria Metelli,Francesco Trovò,Matteo Pirola,Marcello Restelli

from arxiv, Corrected definition of "cumulative increment" (Equation 2) and efficient update (Appendix D)

This paper is in the field of stochastic Multi-Armed Bandits (MABs), i.e., those sequential selection techniques able to learn online using only the feedback given by the chosen option (a.k.a. arm). We study a particular case of the rested and restless bandits in which the arms' expected payoff is monotonically non-decreasing. This characteristic allows designing specifically crafted algorithms that exploit the regularity of the payoffs to provide tight regret bounds. We design an algorithm for the rested case (R-ed-UCB) and one for the restless case (R-less-UCB), providing a regret bound depending on the properties of the instance and, under certain circumstances, of $\widetilde{\mathcal{O}}(T^{\frac{2}{3}})$. We empirically compare our algorithms with state-of-the-art methods for non-stationary MABs over several synthetically generated tasks and an online model selection problem for a real-world dataset. Finally, using synthetic and real-world data, we illustrate the effectiveness of the proposed approaches compared with state-of-the-art algorithms for the non-stationary bandits.

翻译：本文涉及随机多武装强盗(MABs)领域,即仅使用所选选项(a.k.a.a. a. a. arm)提供的反馈,能够在线学习的顺序选择技术。我们研究了一个无休止和无休止的匪徒的个案,在这个个案中,武器预期的回报是单调的,而不是消退。这一特征使得我们能够设计专门设计的算法,利用定期支付来提供严格的遗憾界限。我们设计了一种算法(R-ed-UCB),一种算法(R-less-UCB),一种算法(R-less-UCB),根据实例的特性,提供一定的遗憾。在某些情况下,我们用美元(T ⁇ frac{2 ⁇ 3 ⁇ %%)来说明武器预期的回报是单调的。我们将我们的算法与非固定的MAB公司最先进的方法对几项合成产生的任务和真实世界数据集的在线模式选择问题进行比较。最后,我们使用合成和真实世界的模型数据,用非世界级的算法比较了拟议的系统。

0

相关内容

赌博机/老虎机

赌博机/老虎机

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

专知会员服务

49+阅读 · 2022年11月13日

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

UC.Berkeley CS189讲义教材:《机器学习全面指南》，185页pdf

专知会员服务

162+阅读 · 2020年1月16日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

退化Fisher方程解的渐进性研究

国家自然科学基金

0+阅读 · 2015年12月31日

miR-29a调控PTEN-Akt/Wnt-β-catenin通路促进轴突伸长和神经干细胞增殖修复脊髓损伤的机制

国家自然科学基金

0+阅读 · 2014年12月31日

软骨终板干细胞通过BMP-2介导的Smad依赖性信号通路调节髓核细胞增殖

国家自然科学基金

0+阅读 · 2014年12月31日

DGKε/SNARE信号通路在糖尿病肾病足细胞胰岛素抵抗中的作用及机制

国家自然科学基金

0+阅读 · 2013年12月31日

超声分子成像在体评价突变型低氧诱导因子-1alpha基因在老年鼠缺血下肢血管新生的实验研究

国家自然科学基金

0+阅读 · 2013年12月31日

传染性法氏囊病病毒VP4蛋白与GILZ相互作用抑制I型干扰素表达分子机理的研究

国家自然科学基金

0+阅读 · 2013年12月31日

钙蛋白酶抑制剂对丙烯酰胺神经病的保护作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

新型高稳定全光纤NICE-OHMS色散光谱技术研究

国家自然科学基金

0+阅读 · 2009年12月31日

UGT基因簇进化及调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

家蚕组织蛋白酶D基因表达调控的分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

Projection-free Online Exp-concave Optimization

Arxiv

0+阅读 · 2023年2月9日

Robust and Scalable Bayesian Online Changepoint Detection

Arxiv

0+阅读 · 2023年2月9日

A Constant-per-Iteration Likelihood Ratio Test for Online Changepoint Detection for Exponential Family Models

A Constant-per-Iteration Likelihood Ratio Test for Online Changepoint Detection for Exponential Family Models

Arxiv

0+阅读 · 2023年2月9日

Stochastic Maximum Likelihood Direction Finding in the Presence of Nonuniform Noise Fields

Arxiv

0+阅读 · 2023年2月9日

Optimistic Online Mirror Descent for Bridging Stochastic and Adversarial Online Convex Optimization

Arxiv

0+阅读 · 2023年2月9日

Lazy OCO: Online Convex Optimization on a Switching Budget

Arxiv

0+阅读 · 2023年2月9日

A convolutional neural network of low complexity for tumor anomaly detection

Arxiv

0+阅读 · 2023年2月7日

Towards Understanding the Effects of Evolving the MCTS UCT Selection Policy

Arxiv

0+阅读 · 2023年2月7日

Variance-Aware Sparse Linear Bandits

Arxiv

1+阅读 · 2023年2月6日

Meta-Learning to Cluster

Meta-Learning to Cluster

Arxiv

17+阅读 · 2019年10月30日

VIP会员

文章信息

相关主题

赌博机/老虎机

state-of-the-art

相关VIP内容

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

专知会员服务

49+阅读 · 2022年11月13日

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

UC.Berkeley CS189讲义教材:《机器学习全面指南》，185页pdf

专知会员服务

162+阅读 · 2020年1月16日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《物联网（IoT）中的无人机通信高效控制》135页

《在GNSS信号降级环境中利用共识实现无人机集群稳健协调》

中程单向攻击无人机的战略意义：俄乌战争启示

《面向无人机集群的避障动态传感器覆盖算法》最新38页

相关资讯

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Projection-free Online Exp-concave Optimization

Arxiv

0+阅读 · 2023年2月9日

Robust and Scalable Bayesian Online Changepoint Detection

Arxiv

0+阅读 · 2023年2月9日

A Constant-per-Iteration Likelihood Ratio Test for Online Changepoint Detection for Exponential Family Models

A Constant-per-Iteration Likelihood Ratio Test for Online Changepoint Detection for Exponential Family Models

Arxiv

0+阅读 · 2023年2月9日

Stochastic Maximum Likelihood Direction Finding in the Presence of Nonuniform Noise Fields

Arxiv

0+阅读 · 2023年2月9日

Optimistic Online Mirror Descent for Bridging Stochastic and Adversarial Online Convex Optimization

Arxiv

0+阅读 · 2023年2月9日

Lazy OCO: Online Convex Optimization on a Switching Budget

Arxiv

0+阅读 · 2023年2月9日

A convolutional neural network of low complexity for tumor anomaly detection

Arxiv

0+阅读 · 2023年2月7日

Towards Understanding the Effects of Evolving the MCTS UCT Selection Policy

Arxiv

0+阅读 · 2023年2月7日

Variance-Aware Sparse Linear Bandits

Arxiv

1+阅读 · 2023年2月6日

Meta-Learning to Cluster

Meta-Learning to Cluster

Arxiv

17+阅读 · 2019年10月30日

相关基金

退化Fisher方程解的渐进性研究

国家自然科学基金

0+阅读 · 2015年12月31日

miR-29a调控PTEN-Akt/Wnt-β-catenin通路促进轴突伸长和神经干细胞增殖修复脊髓损伤的机制

国家自然科学基金

0+阅读 · 2014年12月31日

软骨终板干细胞通过BMP-2介导的Smad依赖性信号通路调节髓核细胞增殖

国家自然科学基金

0+阅读 · 2014年12月31日

DGKε/SNARE信号通路在糖尿病肾病足细胞胰岛素抵抗中的作用及机制

国家自然科学基金

0+阅读 · 2013年12月31日

超声分子成像在体评价突变型低氧诱导因子-1alpha基因在老年鼠缺血下肢血管新生的实验研究

国家自然科学基金

0+阅读 · 2013年12月31日

传染性法氏囊病病毒VP4蛋白与GILZ相互作用抑制I型干扰素表达分子机理的研究

国家自然科学基金

0+阅读 · 2013年12月31日

钙蛋白酶抑制剂对丙烯酰胺神经病的保护作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

新型高稳定全光纤NICE-OHMS色散光谱技术研究

国家自然科学基金

0+阅读 · 2009年12月31日

UGT基因簇进化及调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

家蚕组织蛋白酶D基因表达调控的分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员