具有已知制约功能的多能源管理系统的安全强化学习 (Safe reinforcement learning for multi-energy management systems with known constraint functions) - 专知论文

会员服务 ·

0

Learning · 约束 · 泛函 · 可约的 · 优化器 ·

2022 年 7 月 8 日

Safe reinforcement learning for multi-energy management systems with known constraint functions

翻译：具有已知制约功能的多能源管理系统的安全强化学习

Glenn Ceusters,Luis Ramirez Camargo,Rüdiger Franke,Ann Nowé,Maarten Messagie

from arxiv, 25 pages, 12 figures

Reinforcement learning (RL) is a promising optimal control technique for multi-energy management systems. It does not require a model a priori - reducing the upfront and ongoing project-specific engineering effort and is capable of learning better representations of the underlying system dynamics. However, vanilla RL does not provide constraint satisfaction guarantees - resulting in various unsafe interactions within its safety-critical environment. In this paper, we present two novel safe RL methods, namely SafeFallback and GiveSafe, where the safety constraint formulation is decoupled from the RL formulation and which provides hard-constraint satisfaction guarantees both during training (exploration) and exploitation of the (close-to) optimal policy. In a simulated multi-energy systems case study we have shown that both methods start with a significantly higher utility (i.e. useful policy) compared to a vanilla RL benchmark (94,6% and 82,8% compared to 35,5%) and that the proposed SafeFallback method even can outperform the vanilla RL benchmark (102,9% to 100%). We conclude that both methods are viably safety constraint handling techniques capable beyond RL, as demonstrated with random agents while still providing hard-constraint guarantees. Finally, we propose fundamental future work to i.a. improve the constraint functions itself as more data becomes available.

翻译：强化学习(RL)是多能源管理系统中最有希望的最佳控制技术。它不需要先验性的模式 — — 减少前期和正在进行的具体项目工程工作,能够更好地了解基本系统动态的反映。但是,香草RL没有提供约束性满意度保障,导致安全临界环境中的各种不安全互动。在本文中,我们介绍了两种新型的安全安全RL方法,即“安全Fallback”和“GeletSafe”,其中安全限制配方与RL配方脱钩,在培训(勘探)和开发(接近)最佳政策期间提供硬约束性满意度保障。在模拟多能源系统案例研究中,我们发现这两种方法的起点都比香草RL基准(94.6%和82.8%,比35.5%)高得多。拟议的“安全限制配方”方法甚至可以比Vanilla RL基准(102.9 %至100 % ) 。我们的结论是,这两种方法都是在提供超越现有基本安全约束性技术的硬性限制的同时,我们最后提议将安全约束作为基本保证。

0

相关内容

Learning

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

短周期氧化物超晶格中的电荷转移及其应变调控

国家自然科学基金

0+阅读 · 2015年12月31日

FeSe铁基超导薄膜的扫描隧道显微学研究

国家自然科学基金

0+阅读 · 2014年12月31日

赤桉ICE1调控低温胁迫响应的分子机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

功率变换器非线性不稳定行为的washout滤波器控制方法

国家自然科学基金

0+阅读 · 2012年12月31日

SiC低维纳米材料压阻特性及其性能剪裁

国家自然科学基金

0+阅读 · 2012年12月31日

统计模型中的若干组合问题

国家自然科学基金

0+阅读 · 2011年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

ZnO基稀磁半导体薄膜及其ZnO/GaN异质结自旋LED研究

国家自然科学基金

0+阅读 · 2009年12月31日

Active Learning with Effective Scoring Functions for Semi-Supervised Temporal Action Localization

Active Learning with Effective Scoring Functions for Semi-Supervised Temporal Action Localization

Arxiv

0+阅读 · 2022年8月31日

A stabilizing reinforcement learning approach for sampled systems with partially unknown models

Arxiv

0+阅读 · 2022年8月31日

Model-Based Reinforcement Learning with SINDy

Arxiv

0+阅读 · 2022年8月30日

Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation

Arxiv

0+阅读 · 2022年8月30日

A Comparison of Reinforcement Learning Frameworks for Software Testing Tasks

Arxiv

0+阅读 · 2022年8月30日

On stabilizing reinforcement learning without Lyapunov functions

Arxiv

0+阅读 · 2022年8月26日

Prospect Theory-inspired Automated P2P Energy Trading with Q-learning-based Dynamic Pricing

Prospect Theory-inspired Automated P2P Energy Trading with Q-learning-based Dynamic Pricing

Arxiv

0+阅读 · 2022年8月26日

Light-weight probing of unsupervised representations for Reinforcement Learning

Arxiv

0+阅读 · 2022年8月25日

Active Learning for Domain Adaptation: An Energy-based Approach

Arxiv

13+阅读 · 2021年12月2日

A Survey on Reinforcement Learning for Recommender Systems

Arxiv

22+阅读 · 2021年9月22日

VIP会员

文章信息

相关主题

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】以人为中心的强化学习

任务规划与地形分析：现代复杂环境作战导航体系

认知优势：人工智能在国家安全决策中的核心作用

大模型赋能的具身智能：决策与具身学习综述

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Active Learning with Effective Scoring Functions for Semi-Supervised Temporal Action Localization

Active Learning with Effective Scoring Functions for Semi-Supervised Temporal Action Localization

Arxiv

0+阅读 · 2022年8月31日

A stabilizing reinforcement learning approach for sampled systems with partially unknown models

Arxiv

0+阅读 · 2022年8月31日

Model-Based Reinforcement Learning with SINDy

Arxiv

0+阅读 · 2022年8月30日

Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation

Arxiv

0+阅读 · 2022年8月30日

A Comparison of Reinforcement Learning Frameworks for Software Testing Tasks

Arxiv

0+阅读 · 2022年8月30日

On stabilizing reinforcement learning without Lyapunov functions

Arxiv

0+阅读 · 2022年8月26日

Prospect Theory-inspired Automated P2P Energy Trading with Q-learning-based Dynamic Pricing

Prospect Theory-inspired Automated P2P Energy Trading with Q-learning-based Dynamic Pricing

Arxiv

0+阅读 · 2022年8月26日

Light-weight probing of unsupervised representations for Reinforcement Learning

Arxiv

0+阅读 · 2022年8月25日

Active Learning for Domain Adaptation: An Energy-based Approach

Arxiv

13+阅读 · 2021年12月2日

A Survey on Reinforcement Learning for Recommender Systems

Arxiv

22+阅读 · 2021年9月22日

相关基金

短周期氧化物超晶格中的电荷转移及其应变调控

国家自然科学基金

0+阅读 · 2015年12月31日

FeSe铁基超导薄膜的扫描隧道显微学研究

国家自然科学基金

0+阅读 · 2014年12月31日

赤桉ICE1调控低温胁迫响应的分子机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

功率变换器非线性不稳定行为的washout滤波器控制方法

国家自然科学基金

0+阅读 · 2012年12月31日

SiC低维纳米材料压阻特性及其性能剪裁

国家自然科学基金

0+阅读 · 2012年12月31日

统计模型中的若干组合问题

国家自然科学基金

0+阅读 · 2011年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

ZnO基稀磁半导体薄膜及其ZnO/GaN异质结自旋LED研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员