理解 " 初步加强学习 " 中中毒袭击的限度 (Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning) - 专知论文

会员服务 ·

0

Learning · 可理解性 · 强化学习 · 知识 (knowledge) · contrastive ·

2022 年 8 月 29 日

Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning

翻译：理解 " 初步加强学习 " 中中毒袭击的限度

Anshuka Rangi,Haifeng Xu,Long Tran-Thanh,Massimo Franceschetti

from arxiv, Accepted at International Joint Conferences on Artificial Intelligence (IJCAI) 2022

To understand the security threats to reinforcement learning (RL) algorithms, this paper studies poisoning attacks to manipulate \emph{any} order-optimal learning algorithm towards a targeted policy in episodic RL and examines the potential damage of two natural types of poisoning attacks, i.e., the manipulation of \emph{reward} and \emph{action}. We discover that the effect of attacks crucially depend on whether the rewards are bounded or unbounded. In bounded reward settings, we show that only reward manipulation or only action manipulation cannot guarantee a successful attack. However, by combining reward and action manipulation, the adversary can manipulate any order-optimal learning algorithm to follow any targeted policy with $\tilde{\Theta}(\sqrt{T})$ total attack cost, which is order-optimal, without any knowledge of the underlying MDP. In contrast, in unbounded reward settings, we show that reward manipulation attacks are sufficient for an adversary to successfully manipulate any order-optimal learning algorithm to follow any targeted policy using $\tilde{O}(\sqrt{T})$ amount of contamination. Our results reveal useful insights about what can or cannot be achieved by poisoning attacks, and are set to spur more works on the design of robust RL algorithms.

翻译：为了理解对强化学习(RL)算法的安全威胁,本文研究对袭击的毒害性威胁,以操纵 emph{reward} 和\emph{action} 来控制对强化学习(RL) 算法的安全威胁。为了理解对强化学习(RL) 算法的安全威胁,本文研究对袭击的毒害性威胁,以操纵 \ emph{ anny} 秩序优化的学习算法, 以此来对 Associal RLLL(\\ qrt{T}) 的定向政策进行操纵, 并考察两种自然的中毒攻击性攻击, 即操纵\ emph{resward} 和\ emphem{a{ action} 。我们发现, 攻击的效果主要取决于奖赏是否受约束。在受约束的奖赏环境中, 我们显示, 奖赏性攻击的对手足以成功地操纵任何秩序优化学习算法, 来遵循任何目标政策, 使用 $tilde{O} 来操纵任何命令优化的学习算法, 遵循任何目标性的政策。

0

相关内容

Learning

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

MARVELD1基因调控肝细胞癌介入治疗的机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

Copine VII在阿尔茨海默病中的作用机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

执行器故障的大型挠性卫星姿态大角度快速机动容错控制研究

国家自然科学基金

0+阅读 · 2015年12月31日

考虑非定常气动力随机不确定性的气动弹性研究

国家自然科学基金

0+阅读 · 2013年12月31日

Par-4在hTERT非端粒酶活性依赖抗凋亡中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

谷氨酸受体在酒精依赖大鼠冲动性行为中的作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

协助朊病毒感染、致病的lncRNA鉴定及其功能分析

国家自然科学基金

0+阅读 · 2012年12月31日

靶向干预G蛋白偶联受体40对妊娠期糖尿病大鼠胰岛素抵抗及糖稳态的影响

国家自然科学基金

0+阅读 · 2012年12月31日

IGF-2基因印记与PGC-1α转录水平的表观遗传调控在IUGR大鼠胰岛素抵抗的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

适应多类型Insider Attack的入侵检测与精确定位方法的研究

国家自然科学基金

0+阅读 · 2008年12月31日

Boosting Offline Reinforcement Learning via Data Rebalancing

Boosting Offline Reinforcement Learning via Data Rebalancing

Arxiv

0+阅读 · 2022年10月17日

Sample-Efficient Reinforcement Learning of Partially Observable Markov Games

Arxiv

0+阅读 · 2022年10月17日

Breaking the Sample Complexity Barrier to Regret-Optimal Model-Free Reinforcement Learning

Arxiv

0+阅读 · 2022年10月17日

The Impact of Task Underspecification in Evaluating Deep Reinforcement Learning

Arxiv

0+阅读 · 2022年10月16日

Influencing Long-Term Behavior in Multiagent Reinforcement Learning

Arxiv

0+阅读 · 2022年10月15日

Active Exploration for Inverse Reinforcement Learning

Arxiv

0+阅读 · 2022年10月12日

Semi-Supervised Offline Reinforcement Learning with Action-Free Trajectories

Arxiv

0+阅读 · 2022年10月12日

Privacy and Robustness in Federated Learning: Attacks and Defenses

Arxiv

35+阅读 · 2020年12月7日

A Survey on the Explainability of Supervised Machine Learning

Arxiv

24+阅读 · 2020年11月16日

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Arxiv

26+阅读 · 2020年2月10日

VIP会员

文章信息

相关主题

知识 (knowledge)

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《乌克兰无人机产业：志愿者与政策在构建新兴无人机产业中的协同作用》最新报告

《人工智能辅助决策中的数据可视化：系统性综述》

人工智能驱动弹药制造现代化：美国陆军转型之路

《敏捷作战部署中枢纽-辐条基地选址优化研究》80页

相关资讯

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Boosting Offline Reinforcement Learning via Data Rebalancing

Boosting Offline Reinforcement Learning via Data Rebalancing

Arxiv

0+阅读 · 2022年10月17日

Sample-Efficient Reinforcement Learning of Partially Observable Markov Games

Arxiv

0+阅读 · 2022年10月17日

Breaking the Sample Complexity Barrier to Regret-Optimal Model-Free Reinforcement Learning

Arxiv

0+阅读 · 2022年10月17日

The Impact of Task Underspecification in Evaluating Deep Reinforcement Learning

Arxiv

0+阅读 · 2022年10月16日

Influencing Long-Term Behavior in Multiagent Reinforcement Learning

Arxiv

0+阅读 · 2022年10月15日

Active Exploration for Inverse Reinforcement Learning

Arxiv

0+阅读 · 2022年10月12日

Semi-Supervised Offline Reinforcement Learning with Action-Free Trajectories

Arxiv

0+阅读 · 2022年10月12日

Privacy and Robustness in Federated Learning: Attacks and Defenses

Arxiv

35+阅读 · 2020年12月7日

A Survey on the Explainability of Supervised Machine Learning

Arxiv

24+阅读 · 2020年11月16日

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Arxiv

26+阅读 · 2020年2月10日

相关基金

MARVELD1基因调控肝细胞癌介入治疗的机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

Copine VII在阿尔茨海默病中的作用机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

执行器故障的大型挠性卫星姿态大角度快速机动容错控制研究

国家自然科学基金

0+阅读 · 2015年12月31日

考虑非定常气动力随机不确定性的气动弹性研究

国家自然科学基金

0+阅读 · 2013年12月31日

Par-4在hTERT非端粒酶活性依赖抗凋亡中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

谷氨酸受体在酒精依赖大鼠冲动性行为中的作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

协助朊病毒感染、致病的lncRNA鉴定及其功能分析

国家自然科学基金

0+阅读 · 2012年12月31日

靶向干预G蛋白偶联受体40对妊娠期糖尿病大鼠胰岛素抵抗及糖稳态的影响

国家自然科学基金

0+阅读 · 2012年12月31日

IGF-2基因印记与PGC-1α转录水平的表观遗传调控在IUGR大鼠胰岛素抵抗的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

适应多类型Insider Attack的入侵检测与精确定位方法的研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员