有奖励足以证明其手段吗？在MACHIAVELLI基准测试中度量奖励与道德行为之间的权衡 (Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark) - 专知论文

会员服务 ·

0

基准测试 · 情境 · 基准 · 度量 · 语言模型 ·

2023 年 5 月 1 日

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

翻译：有奖励足以证明其手段吗？在MACHIAVELLI基准测试中度量奖励与道德行为之间的权衡

Alexander Pan,Chan Jun Shern,Andy Zou,Nathaniel Li,Steven Basart,Thomas Woodside,Jonathan Ng,Hanlin Zhang,Scott Emmons,Dan Hendrycks

from arxiv, ICML 2023 Oral; 31 pages, 5 figures

Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity. So do agents naturally learn to be Machiavellian? And how do we measure these behaviors in general-purpose models such as GPT-4? Towards answering these questions, we introduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure games containing over half a million rich, diverse scenarios that center on social decision-making. Scenario labeling is automated with LMs, which are more performant than human annotators. We mathematize dozens of harmful behaviors and use our annotations to evaluate agents' tendencies to be power-seeking, cause disutility, and commit ethical violations. We observe some tension between maximizing reward and behaving ethically. To improve this trade-off, we investigate LM-based methods to steer agents' towards less harmful behaviors. Our results show that agents can both act competently and morally, so concrete progress can currently be made in machine ethics--designing agents that are Pareto improvements in both safety and capabilities.

翻译：传统上，人工智能代理的培训方法是最大化奖励，这可能会激励追求权力和欺骗，类似于语言模型中的下一个令牌预测可能会激励毒性。那么代理是否自然而然地学会了马基雅维利主义？我们如何在通用模型（如GPT-4）中衡量这些行为？为了回答这些问题，我们引入了MACHIAVELLI，一个包含超过50万个围绕社会决策制定的丰富多样情境的134款冒险文艺游戏的基准测试。情境标注由比人类注释者更高性能的语言模型自动执行。我们数学化了数十种有害行为，并使用我们的注释来评估代理对于寻求权力和造成不适、违反道德的倾向。我们观察到，在最大化奖励和行使道德之间存在一定的紧张关系。为了改善这种权衡，我们研究了基于语言模型的方法，以引导代理朝向更少有害的行为方向。我们的研究结果表明，代理可以既能表现得有能力又能表现得道德，因此我们可以在机器伦理学方面取得实质性的进展，即设计出既安全又具备能力的Pareto改进型代理。

0

相关内容

基准测试

基准测试是指通过设计科学的测试方法、测试工具和测试系统，实现对一类测试对象的某项性能指标进行定量的和可对比的测试。

GPT-4等大模型懂因果么？ Meta等最新《大型语言模型能从相关性中推断因果关系吗》17种LLM表现一般，GPT-4也不行

GPT-4等大模型懂因果么？ Meta等最新《大型语言模型能从相关性中推断因果关系吗》17种LLM表现一般，GPT-4也不行

专知会员服务

60+阅读 · 2023年6月12日

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

专知会员服务

49+阅读 · 2022年11月13日

《JADC2 Update—— The What to the How》美国国防信息系统局（DISA）10页slides

《JADC2 Update—— The What to the How》美国国防信息系统局（DISA）10页slides

专知会员服务

49+阅读 · 2022年6月8日

语言视觉预训练语言模型揭密，Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

语言视觉预训练语言模型揭密，Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

专知会员服务

36+阅读 · 2020年5月20日

【斯坦福大学-PNAS2020】人工智能中深度学习的不合理有效性unreasonable effectiveness of DL

专知会员服务

14+阅读 · 2020年2月23日

【斯坦福大学AAAI2020】跨越因果层次的概率推理，Probabilistic Reasoning across the Causal Hierarchy

【斯坦福大学AAAI2020】跨越因果层次的概率推理，Probabilistic Reasoning across the Causal Hierarchy

专知会员服务

46+阅读 · 2020年1月11日

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

专知会员服务

16+阅读 · 2019年12月10日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

GNN 新基准！Long Range Graph Benchmark

GNN 新基准！Long Range Graph Benchmark

图与推荐

0+阅读 · 2022年10月18日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

【论文推荐】最新七篇视觉问答（VQA）相关论文—差别注意力机制、视觉问题推理、视觉对话、数据可视化、记忆增强网络、显式推理

【论文推荐】最新七篇视觉问答（VQA）相关论文—差别注意力机制、视觉问题推理、视觉对话、数据可视化、记忆增强网络、显式推理

专知

17+阅读 · 2018年4月19日

强化学习初探 - 从多臂老虎机问题说起

强化学习初探 - 从多臂老虎机问题说起

专知

10+阅读 · 2018年4月3日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

吸入氢气调节BMPR2表达、防治慢阻肺的作用和机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

ARVCF调节cadherin/catenin复合体介导的细胞间黏附的分子机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

MDM2/E2F1通过调控RARa蛋白水平影响骨肉瘤细胞分化的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于风险能量的BT债务规模边界及风险控制有效性研究

国家自然科学基金

0+阅读 · 2014年12月31日

PKCα与UNC5B相互作用调控膀胱癌细胞药物敏感性的分子机制

国家自然科学基金

0+阅读 · 2013年12月31日

网络媒体影响政治态度的因果机制的实验研究

国家自然科学基金

3+阅读 · 2012年12月31日

小世界模型的伪随机性质

国家自然科学基金

0+阅读 · 2011年12月31日

E3泛素连接酶CHIP在前列腺癌雄激素非依赖性形成中的作用和机制

国家自然科学基金

0+阅读 · 2011年12月31日

群集行为控制的理论与实验研究

国家自然科学基金

1+阅读 · 2009年12月31日

分数布朗运动环境下金融保险中优化问题的研究

国家自然科学基金

0+阅读 · 2009年12月31日

Trained Transformers Learn Linear Models In-Context

Arxiv

1+阅读 · 2023年6月16日

Are ChatGPT and Other Similar Systems the Modern Lernaean Hydras of AI?

Arxiv

0+阅读 · 2023年6月15日

Unprocessing Seven Years of Algorithmic Fairness

Arxiv

0+阅读 · 2023年6月15日

Reward-Free Curricula for Training Robust World Models

Arxiv

0+阅读 · 2023年6月15日

Evolutionary Curriculum Training for DRL-Based Navigation Systems

Arxiv

0+阅读 · 2023年6月15日

Improving Reading Comprehension Question Generation with Data Augmentation and Overgenerate-and-rank

Arxiv

0+阅读 · 2023年6月15日

Analysis and Comparison of Classification Metrics

Arxiv

0+阅读 · 2023年6月14日

Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

Arxiv

12+阅读 · 2023年4月26日

Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

Arxiv

11+阅读 · 2021年9月3日

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Arxiv

20+阅读 · 2020年3月10日

VIP会员

文章信息

相关主题

相关VIP内容

GPT-4等大模型懂因果么？ Meta等最新《大型语言模型能从相关性中推断因果关系吗》17种LLM表现一般，GPT-4也不行

GPT-4等大模型懂因果么？ Meta等最新《大型语言模型能从相关性中推断因果关系吗》17种LLM表现一般，GPT-4也不行

专知会员服务

60+阅读 · 2023年6月12日

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

专知会员服务

49+阅读 · 2022年11月13日

《JADC2 Update—— The What to the How》美国国防信息系统局（DISA）10页slides

《JADC2 Update—— The What to the How》美国国防信息系统局（DISA）10页slides

专知会员服务

49+阅读 · 2022年6月8日

语言视觉预训练语言模型揭密，Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

语言视觉预训练语言模型揭密，Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

专知会员服务

36+阅读 · 2020年5月20日

【斯坦福大学-PNAS2020】人工智能中深度学习的不合理有效性unreasonable effectiveness of DL

专知会员服务

14+阅读 · 2020年2月23日

【斯坦福大学AAAI2020】跨越因果层次的概率推理，Probabilistic Reasoning across the Causal Hierarchy

【斯坦福大学AAAI2020】跨越因果层次的概率推理，Probabilistic Reasoning across the Causal Hierarchy

专知会员服务

46+阅读 · 2020年1月11日

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

专知会员服务

16+阅读 · 2019年12月10日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

自动驾驶轨迹规划中的基础模型：进展综述与开放挑战

《用于提升多域战备的大型语言模型辅助场景生成器》报告

【斯坦福博士论文】为人类使用优化 AI 模型

国防领域人工智能规模化应用的理论与实践

相关资讯

GNN 新基准！Long Range Graph Benchmark

GNN 新基准！Long Range Graph Benchmark

图与推荐

0+阅读 · 2022年10月18日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

【论文推荐】最新七篇视觉问答（VQA）相关论文—差别注意力机制、视觉问题推理、视觉对话、数据可视化、记忆增强网络、显式推理

【论文推荐】最新七篇视觉问答（VQA）相关论文—差别注意力机制、视觉问题推理、视觉对话、数据可视化、记忆增强网络、显式推理

专知

17+阅读 · 2018年4月19日

强化学习初探 - 从多臂老虎机问题说起

强化学习初探 - 从多臂老虎机问题说起

专知

10+阅读 · 2018年4月3日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

相关论文

Trained Transformers Learn Linear Models In-Context

Arxiv

1+阅读 · 2023年6月16日

Are ChatGPT and Other Similar Systems the Modern Lernaean Hydras of AI?

Arxiv

0+阅读 · 2023年6月15日

Unprocessing Seven Years of Algorithmic Fairness

Arxiv

0+阅读 · 2023年6月15日

Reward-Free Curricula for Training Robust World Models

Arxiv

0+阅读 · 2023年6月15日

Evolutionary Curriculum Training for DRL-Based Navigation Systems

Arxiv

0+阅读 · 2023年6月15日

Improving Reading Comprehension Question Generation with Data Augmentation and Overgenerate-and-rank

Arxiv

0+阅读 · 2023年6月15日

Analysis and Comparison of Classification Metrics

Arxiv

0+阅读 · 2023年6月14日

Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

Arxiv

12+阅读 · 2023年4月26日

Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

Arxiv

11+阅读 · 2021年9月3日

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Arxiv

20+阅读 · 2020年3月10日

相关基金

吸入氢气调节BMPR2表达、防治慢阻肺的作用和机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

ARVCF调节cadherin/catenin复合体介导的细胞间黏附的分子机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

MDM2/E2F1通过调控RARa蛋白水平影响骨肉瘤细胞分化的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于风险能量的BT债务规模边界及风险控制有效性研究

国家自然科学基金

0+阅读 · 2014年12月31日

PKCα与UNC5B相互作用调控膀胱癌细胞药物敏感性的分子机制

国家自然科学基金

0+阅读 · 2013年12月31日

网络媒体影响政治态度的因果机制的实验研究

国家自然科学基金

3+阅读 · 2012年12月31日

小世界模型的伪随机性质

国家自然科学基金

0+阅读 · 2011年12月31日

E3泛素连接酶CHIP在前列腺癌雄激素非依赖性形成中的作用和机制

国家自然科学基金

0+阅读 · 2011年12月31日

群集行为控制的理论与实验研究

国家自然科学基金

1+阅读 · 2009年12月31日

分数布朗运动环境下金融保险中优化问题的研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员