以RL为基础的政策优化方法以适应性稳定认证为指导 (A RL-based Policy Optimization Method Guided by Adaptive Stability Certification) - 专知论文

会员服务 ·

0

Lyapunov · 优化器 · 约束 · Learning · 最优化 ·

2023 年 1 月 2 日

A RL-based Policy Optimization Method Guided by Adaptive Stability Certification

翻译：以RL为基础的政策优化方法以适应性稳定认证为指导

Shengjie Wang,Fengbo Lan,Xiang Zheng,Yuxue Cao,Oluwatosin Oseni,Haotian Xu,Yang Gao,Tao Zhang

from arxiv, 27 pages, 11 figues

In contrast to the control-theoretic methods, the lack of stability guarantee remains a significant problem for model-free reinforcement learning (RL) methods. Jointly learning a policy and a Lyapunov function has recently become a promising approach to ensuring the whole system with a stability guarantee. However, the classical Lyapunov constraints researchers introduced cannot stabilize the system during the sampling-based optimization. Therefore, we propose the Adaptive Stability Certification (ASC), making the system reach sampling-based stability. Because the ASC condition can search for the optimal policy heuristically, we design the Adaptive Lyapunov-based Actor-Critic (ALAC) algorithm based on the ASC condition. Meanwhile, our algorithm avoids the optimization problem that a variety of constraints are coupled into the objective in current approaches. When evaluated on ten robotic tasks, our method achieves lower accumulated cost and fewer stability constraint violations than previous studies.

翻译：与控制理论方法不同,缺乏稳定性保障仍然是无模型强化学习方法的重大问题。共同学习政策和Lyapunov函数最近成为确保整个系统有稳定保障的一个很有希望的方法。然而,古典的Lyapunov限制研究者在取样优化期间无法稳定系统。因此,我们建议采用适应性稳定认证,使系统达到基于取样的稳定。由于ASC条件可以超常地寻找最佳政策,我们根据ASC条件设计了适应性Lyapunov-ALAC(ALAC)算法。与此同时,我们的算法避免了优化问题,即各种限制与当前方法的目标相结合。在对10项机器人任务进行评估时,我们的方法的累积成本较低,稳定性限制比以往的研究要少。

0

相关内容

Lyapunov

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

零样本文本分类，Zero-Shot Learning for Text Classification

零样本文本分类，Zero-Shot Learning for Text Classification

专知会员服务

97+阅读 · 2020年5月31日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

Tet3介导的表观遗传修饰对CD4+T细胞分化及Th17细胞功能的调控

国家自然科学基金

0+阅读 · 2016年12月31日

基于自主学习的Ad hoc Agent序贯决策研究

国家自然科学基金

44+阅读 · 2015年12月31日

钢管再生混凝土混合结构体系的抗震性能及灾变控制研究

国家自然科学基金

0+阅读 · 2014年12月31日

血管壁微环境Oxysterol/Ox-LDL通过氧化固醇结合蛋白ORP4L诱导巨噬细胞凋亡的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

血小板介导肿瘤血管形成和稳定性的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

内皮细胞TRPV4-SKCa3耦联稳态失调在高血压血管功能稳态失调中的作用及机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

β防御素在牙龈卟啉单胞菌脂多糖促动脉粥样硬化形成中的保护作用

国家自然科学基金

0+阅读 · 2012年12月31日

实时安全关键系统的建模、仿真与验证

国家自然科学基金

1+阅读 · 2012年12月31日

树枝状分子功能性液晶凝胶的制备、性质与应用研究

国家自然科学基金

0+阅读 · 2009年12月31日

蛋白激酶A对血小板生理功能的调控作用及其机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

AdaSAM: Boosting Sharpness-Aware Minimization with Adaptive Learning Rate and Momentum for Training Deep Neural Networks

Arxiv

0+阅读 · 2023年3月1日

Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Arxiv

0+阅读 · 2023年3月1日

Ensemble-based gradient inference for particle methods in optimization and sampling

Arxiv

0+阅读 · 2023年3月1日

Learning to Control Autonomous Fleets from Observation via Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年2月28日

Backstepping Temporal Difference Learning

Arxiv

0+阅读 · 2023年2月28日

SafeLight: A Reinforcement Learning Method toward Collision-free Traffic Signal Control

Arxiv

0+阅读 · 2023年2月27日

RTAW: An Attention Inspired Reinforcement Learning Method for Multi-Robot Task Allocation in Warehouse Environments

RTAW: An Attention Inspired Reinforcement Learning Method for Multi-Robot Task Allocation in Warehouse Environments

Arxiv

0+阅读 · 2023年2月27日

The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年2月27日

When Source-Free Domain Adaptation Meets Learning with Noisy Labels

Arxiv

0+阅读 · 2023年2月24日

High-Speed VLSI Architectures for Modular Polynomial Multiplication via Fast Filtering and Applications to Lattice-Based Cryptography

Arxiv

0+阅读 · 2023年2月24日

VIP会员

文章信息

相关主题

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

零样本文本分类，Zero-Shot Learning for Text Classification

零样本文本分类，Zero-Shot Learning for Text Classification

专知会员服务

97+阅读 · 2020年5月31日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《人与智能体在系统工程建模语言V2任务中的性能表现：基于用户中心化的评估方法》308页

《数据安全国家标准体系（2025版）》征求意见稿

AlphaMosaic：人工智能赋能的作战管理系统

《军事行动中通信平台的战略价值：提升战术效能与作战优势》

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

AdaSAM: Boosting Sharpness-Aware Minimization with Adaptive Learning Rate and Momentum for Training Deep Neural Networks

Arxiv

0+阅读 · 2023年3月1日

Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Arxiv

0+阅读 · 2023年3月1日

Ensemble-based gradient inference for particle methods in optimization and sampling

Arxiv

0+阅读 · 2023年3月1日

Learning to Control Autonomous Fleets from Observation via Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年2月28日

Backstepping Temporal Difference Learning

Arxiv

0+阅读 · 2023年2月28日

SafeLight: A Reinforcement Learning Method toward Collision-free Traffic Signal Control

Arxiv

0+阅读 · 2023年2月27日

RTAW: An Attention Inspired Reinforcement Learning Method for Multi-Robot Task Allocation in Warehouse Environments

RTAW: An Attention Inspired Reinforcement Learning Method for Multi-Robot Task Allocation in Warehouse Environments

Arxiv

0+阅读 · 2023年2月27日

The Provable Benefits of Unsupervised Data Sharing for Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年2月27日

When Source-Free Domain Adaptation Meets Learning with Noisy Labels

Arxiv

0+阅读 · 2023年2月24日

High-Speed VLSI Architectures for Modular Polynomial Multiplication via Fast Filtering and Applications to Lattice-Based Cryptography

Arxiv

0+阅读 · 2023年2月24日

相关基金

Tet3介导的表观遗传修饰对CD4+T细胞分化及Th17细胞功能的调控

国家自然科学基金

0+阅读 · 2016年12月31日

基于自主学习的Ad hoc Agent序贯决策研究

国家自然科学基金

44+阅读 · 2015年12月31日

钢管再生混凝土混合结构体系的抗震性能及灾变控制研究

国家自然科学基金

0+阅读 · 2014年12月31日

血管壁微环境Oxysterol/Ox-LDL通过氧化固醇结合蛋白ORP4L诱导巨噬细胞凋亡的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

血小板介导肿瘤血管形成和稳定性的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

内皮细胞TRPV4-SKCa3耦联稳态失调在高血压血管功能稳态失调中的作用及机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

β防御素在牙龈卟啉单胞菌脂多糖促动脉粥样硬化形成中的保护作用

国家自然科学基金

0+阅读 · 2012年12月31日

实时安全关键系统的建模、仿真与验证

国家自然科学基金

1+阅读 · 2012年12月31日

树枝状分子功能性液晶凝胶的制备、性质与应用研究

国家自然科学基金

0+阅读 · 2009年12月31日

蛋白激酶A对血小板生理功能的调控作用及其机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员