战略操纵下的组合强盗 (Combinatorial Bandits under Strategic Manipulations) - 专知论文

会员服务 ·

0

赌博机/老虎机 · Extensibility · ARM · 在线 · 稳健性 ·

2021 年 11 月 19 日

Combinatorial Bandits under Strategic Manipulations

翻译：战略操纵下的组合强盗

Jing Dong,Ke Li,Shuai Li,Baoxiang Wang

Strategic behavior against sequential learning methods, such as "click framing" in real recommendation systems, have been widely observed. Motivated by such behavior we study the problem of combinatorial multi-armed bandits (CMAB) under strategic manipulations of rewards, where each arm can modify the emitted reward signals for its own interest. This characterization of the adversarial behavior is a relaxation of previously well-studied settings such as adversarial attacks and adversarial corruption. We propose a strategic variant of the combinatorial UCB algorithm, which has a regret of at most $O(m\log T + m B_{max})$ under strategic manipulations, where $T$ is the time horizon, $m$ is the number of arms, and $B_{max}$ is the maximum budget of an arm. We provide lower bounds on the budget for arms to incur certain regret of the bandit algorithm. Extensive experiments on online worker selection for crowdsourcing systems, online influence maximization and online recommendations with both synthetic and real datasets corroborate our theoretical findings on robustness and regret bounds, in a variety of regimes of manipulation budgets.

翻译：对抗性行为的特征是放松了先前研究周密的环境,如对抗性攻击和对抗性腐败。我们提出了一个组合式UCB算法的战略变体,该算法对在战略操纵下花费最多为$O(m\log T+m B ⁇ max})的负数($O(m\log T+m B ⁇ max})的负数($T)的负数($T)是时间范围,$m是武器的数量,$B ⁇ max}是武器的最大预算。我们提供了较低的武器预算约束,以引起强盗算法的某些遗憾。我们用合成和真实数据集对在线工人选择众包系统、在线影响最大化和在线建议进行了广泛的实验,证实了我们在各种操纵预算制度中对稳健和遗憾界限的理论结论。

0

相关内容

赌博机/老虎机

赌博机/老虎机

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

ICLR2021 | 初探GNN的表示能力

专知会员服务

28+阅读 · 2021年5月2日

ICLR2021放榜了！ 687篇入选34篇得满分！ 48篇orals，108篇spotlights，531篇poster

ICLR2021放榜了！ 687篇入选34篇得满分！ 48篇orals，108篇spotlights，531篇poster

专知会员服务

24+阅读 · 2021年1月13日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【KDD2020-清华大学】理解图表示学习中的负采样，Understanding Negative Sampling

【KDD2020-清华大学】理解图表示学习中的负采样，Understanding Negative Sampling

专知会员服务

63+阅读 · 2020年5月23日

【硬核书】数学博弈论与应用，431页pdf，Mathematical Game Theory and Applications

【硬核书】数学博弈论与应用，431页pdf，Mathematical Game Theory and Applications

专知会员服务

170+阅读 · 2020年4月18日

【CCL 2019】ATT-第19期：Frontiers in Network Embedding and GCN （崔鹏）

【CCL 2019】ATT-第19期：Frontiers in Network Embedding and GCN （崔鹏）

专知会员服务

44+阅读 · 2019年11月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

RL 真经

CreateAMind

5+阅读 · 2018年12月28日

Disentangled的假设的探讨

Disentangled的假设的探讨

CreateAMind

9+阅读 · 2018年12月10日

【学界】AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

【学界】AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

GAN生成式对抗网络

8+阅读 · 2018年11月4日

AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

专知

11+阅读 · 2018年11月4日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

Minimax Demographic Group Fairness in Federated Learning

Minimax Demographic Group Fairness in Federated Learning

Arxiv

0+阅读 · 2022年1月25日

Almost Optimal Variance-Constrained Best Arm Identification

Arxiv

0+阅读 · 2022年1月25日

Weak Signal Inclusion Under Sparsity and Dependence

Arxiv

0+阅读 · 2022年1月24日

VCG Mechanism Design with Unknown Agent Values under Stochastic Bandit Feedback

Arxiv

0+阅读 · 2022年1月23日

A Lyapunov-Based Methodology for Constrained Optimization with Bandit Feedback

Arxiv

0+阅读 · 2022年1月23日

An Improved Lower Bound for Multi-Access Coded Caching

Arxiv

0+阅读 · 2022年1月22日

Noisy linear inverse problems under convex constraints: Exact risk asymptotics in high dimensions

Arxiv

0+阅读 · 2022年1月20日

Policy Gradient Bayesian Robust Optimization for Imitation Learning

Arxiv

5+阅读 · 2021年6月11日

Attribute-Guided Adversarial Training for Robustness to Natural Perturbations

Arxiv

15+阅读 · 2020年12月3日

Deep Reinforcement Learning for List-wise Recommendations

Arxiv

13+阅读 · 2018年1月5日

VIP会员

文章信息

相关主题

赌博机/老虎机

相关VIP内容

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

ICLR2021 | 初探GNN的表示能力

专知会员服务

28+阅读 · 2021年5月2日

ICLR2021放榜了！ 687篇入选34篇得满分！ 48篇orals，108篇spotlights，531篇poster

ICLR2021放榜了！ 687篇入选34篇得满分！ 48篇orals，108篇spotlights，531篇poster

专知会员服务

24+阅读 · 2021年1月13日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【KDD2020-清华大学】理解图表示学习中的负采样，Understanding Negative Sampling

【KDD2020-清华大学】理解图表示学习中的负采样，Understanding Negative Sampling

专知会员服务

63+阅读 · 2020年5月23日

【硬核书】数学博弈论与应用，431页pdf，Mathematical Game Theory and Applications

【硬核书】数学博弈论与应用，431页pdf，Mathematical Game Theory and Applications

专知会员服务

170+阅读 · 2020年4月18日

【CCL 2019】ATT-第19期：Frontiers in Network Embedding and GCN （崔鹏）

【CCL 2019】ATT-第19期：Frontiers in Network Embedding and GCN （崔鹏）

专知会员服务

44+阅读 · 2019年11月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

基于大型语言模型的网络威胁情报：利用LLM提取MITRE ATT&CK技术 | 最新文献

无人机（UAV）战略：区域大国与暴力非国家行为体在中东冲突中对无人机的运用 | 130页

神经技术与未来无人机战争的交汇点 | 最新报告

美国从“蛛网行动”中汲取轰炸机舰队保护教训

相关资讯

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

RL 真经

CreateAMind

5+阅读 · 2018年12月28日

Disentangled的假设的探讨

Disentangled的假设的探讨

CreateAMind

9+阅读 · 2018年12月10日

【学界】AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

【学界】AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

GAN生成式对抗网络

8+阅读 · 2018年11月4日

AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

AAAI2019论文抢鲜看！48篇自然语言处理/计算机视觉/机器学习最新接受论文！

专知

11+阅读 · 2018年11月4日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Minimax Demographic Group Fairness in Federated Learning

Minimax Demographic Group Fairness in Federated Learning

Arxiv

0+阅读 · 2022年1月25日

Almost Optimal Variance-Constrained Best Arm Identification

Arxiv

0+阅读 · 2022年1月25日

Weak Signal Inclusion Under Sparsity and Dependence

Arxiv

0+阅读 · 2022年1月24日

VCG Mechanism Design with Unknown Agent Values under Stochastic Bandit Feedback

Arxiv

0+阅读 · 2022年1月23日

A Lyapunov-Based Methodology for Constrained Optimization with Bandit Feedback

Arxiv

0+阅读 · 2022年1月23日

An Improved Lower Bound for Multi-Access Coded Caching

Arxiv

0+阅读 · 2022年1月22日

Noisy linear inverse problems under convex constraints: Exact risk asymptotics in high dimensions

Arxiv

0+阅读 · 2022年1月20日

Policy Gradient Bayesian Robust Optimization for Imitation Learning

Arxiv

5+阅读 · 2021年6月11日

Attribute-Guided Adversarial Training for Robustness to Natural Perturbations

Arxiv

15+阅读 · 2020年12月3日

Deep Reinforcement Learning for List-wise Recommendations

Arxiv

13+阅读 · 2018年1月5日

微信扫码咨询专知VIP会员