受资源制约的目标(POMDPs) (Shielding in Resource-Constrained Goal POMDPs) - 专知论文

会员服务 ·

0

Agent · 部分可观测马尔可夫决策过程 · Markov · 最优化 · Processing（编程语言） ·

2022 年 11 月 28 日

Shielding in Resource-Constrained Goal POMDPs

翻译：受资源制约的目标(POMDPs)

Michal Ajdarów,Šimon Brlej,Petr Novotný

We consider partially observable Markov decision processes (POMDPs) modeling an agent that needs a supply of a certain resource (e.g., electricity stored in batteries) to operate correctly. The resource is consumed by agent's actions and can be replenished only in certain states. The agent aims to minimize the expected cost of reaching some goal while preventing resource exhaustion, a problem we call \emph{resource-constrained goal optimization} (RSGO). We take a two-step approach to the RSGO problem. First, using formal methods techniques, we design an algorithm computing a \emph{shield} for a given scenario: a procedure that observes the agent and prevents it from using actions that might eventually lead to resource exhaustion. Second, we augment the POMCP heuristic search algorithm for POMDP planning with our shields to obtain an algorithm solving the RSGO problem. We implement our algorithm and present experiments showing its applicability to benchmarks from the literature.

翻译：我们认为,部分可见的Markov决策程序(POMDPs)可以模拟需要提供某种资源(例如电池中储存的电力)才能正确运作的代理商。该资源被代理商的行动消耗,只能在某些州得到补充。该代理商的目的是在防止资源耗竭的同时尽可能降低达到某种目标的预期成本,而防止资源耗竭,我们称之为“资源受资源限制的目标优化”的问题。我们对RSGO问题采取了分两步走的办法。首先,使用正规方法技术,我们设计了一种算法,计算出某种特定情景:一种观察该代理商的程序,防止其使用最终可能导致资源耗竭的行动。第二,我们用防护罩加强POMCP规划POMDP的超速搜索算法,以获得解决RSGO问题的算法。我们实施了我们的算法,并提出了实验,表明其适用于文献基准。

0

相关内容

Agent

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

80+阅读 · 2020年7月26日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

基于情景模拟与兵棋推演的城市暴雨内涝灾害居民应急避难研究

国家自然科学基金

3+阅读 · 2015年12月31日

偕二氟取代Combretastatins衍生物的设计与合成

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

20世纪50年代以来青藏高原气温变化的不确定性定量评估

国家自然科学基金

1+阅读 · 2013年12月31日

全气候作用下沥青混合料中沥青纳米级老化机理和老化动力学研究

国家自然科学基金

0+阅读 · 2012年12月31日

POLD1基因的癌性表达及其在乳腺癌中对细胞恶性表型的影响

国家自然科学基金

0+阅读 · 2012年12月31日

PC-1和CHIP通过蛋白酶体途径共同调控M期雄激素受体降解

国家自然科学基金

0+阅读 · 2012年12月31日

淫羊藿素（ICT）通过激动AhR降解AR的作用抑制前列腺癌的研究

国家自然科学基金

0+阅读 · 2011年12月31日

船舶纵摇与垂荡减摇新方法及水动力机理研究

国家自然科学基金

1+阅读 · 2011年12月31日

磁性Pickering乳液界面流变学研究

国家自然科学基金

0+阅读 · 2008年12月31日

Singularity-aware Reinforcement Learning

Arxiv

0+阅读 · 2023年1月30日

Sample Efficient Deep Reinforcement Learning via Local Planning

Arxiv

0+阅读 · 2023年1月29日

Streaming LifeLong Learning With Any-Time Inference

Arxiv

0+阅读 · 2023年1月27日

A Deep Learning Method for Comparing Bayesian Hierarchical Models

Arxiv

0+阅读 · 2023年1月27日

Constrained Parameter Inference as a Principle for Learning

Arxiv

0+阅读 · 2023年1月27日

Provably Efficient Causal Model-Based Reinforcement Learning for Systematic Generalization

Arxiv

0+阅读 · 2023年1月27日

Demystifying Reinforcement Learning in Time-Varying Systems

Arxiv

0+阅读 · 2023年1月26日

Continual Depth-limited Responses for Computing Counter-strategies in Sequential Games

Arxiv

0+阅读 · 2023年1月26日

Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

Arxiv

11+阅读 · 2021年9月3日

Curriculum Learning: A Survey

Arxiv

24+阅读 · 2021年1月25日

VIP会员

文章信息

相关主题

部分可观测马尔可夫决策过程

Processing（编程语言）

相关VIP内容

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

80+阅读 · 2020年7月26日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

卫星导航技术发展综述

《美军"僚机"联合能力技术演示项目：有人-无人火炮作战》41页报告

美军条令《火力指挥》116页

可解释的人工智能在生物医学图像分析中的应用综述

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Singularity-aware Reinforcement Learning

Arxiv

0+阅读 · 2023年1月30日

Sample Efficient Deep Reinforcement Learning via Local Planning

Arxiv

0+阅读 · 2023年1月29日

Streaming LifeLong Learning With Any-Time Inference

Arxiv

0+阅读 · 2023年1月27日

A Deep Learning Method for Comparing Bayesian Hierarchical Models

Arxiv

0+阅读 · 2023年1月27日

Constrained Parameter Inference as a Principle for Learning

Arxiv

0+阅读 · 2023年1月27日

Provably Efficient Causal Model-Based Reinforcement Learning for Systematic Generalization

Arxiv

0+阅读 · 2023年1月27日

Demystifying Reinforcement Learning in Time-Varying Systems

Arxiv

0+阅读 · 2023年1月26日

Continual Depth-limited Responses for Computing Counter-strategies in Sequential Games

Arxiv

0+阅读 · 2023年1月26日

Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

Arxiv

11+阅读 · 2021年9月3日

Curriculum Learning: A Survey

Arxiv

24+阅读 · 2021年1月25日

相关基金

基于情景模拟与兵棋推演的城市暴雨内涝灾害居民应急避难研究

国家自然科学基金

3+阅读 · 2015年12月31日

偕二氟取代Combretastatins衍生物的设计与合成

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

20世纪50年代以来青藏高原气温变化的不确定性定量评估

国家自然科学基金

1+阅读 · 2013年12月31日

全气候作用下沥青混合料中沥青纳米级老化机理和老化动力学研究

国家自然科学基金

0+阅读 · 2012年12月31日

POLD1基因的癌性表达及其在乳腺癌中对细胞恶性表型的影响

国家自然科学基金

0+阅读 · 2012年12月31日

PC-1和CHIP通过蛋白酶体途径共同调控M期雄激素受体降解

国家自然科学基金

0+阅读 · 2012年12月31日

淫羊藿素（ICT）通过激动AhR降解AR的作用抑制前列腺癌的研究

国家自然科学基金

0+阅读 · 2011年12月31日

船舶纵摇与垂荡减摇新方法及水动力机理研究

国家自然科学基金

1+阅读 · 2011年12月31日

磁性Pickering乳液界面流变学研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员