凸优化基于策略适应的分布转移补偿 (Convex Optimization-based Policy Adaptation to Compensate for Distributional Shifts) - 专知论文

会员服务 ·

0

分布转移 · 最优 · 补偿 · 系统 · 控制器 ·

2023 年 4 月 5 日

Convex Optimization-based Policy Adaptation to Compensate for Distributional Shifts

翻译：凸优化基于策略适应的分布转移补偿

Navid Hashemi,Justin Ruths,Jyotirmoy V. Deshmukh

Many real-world systems often involve physical components or operating environments with highly nonlinear and uncertain dynamics. A number of different control algorithms can be used to design optimal controllers for such systems, assuming a reasonably high-fidelity model of the actual system. However, the assumptions made on the stochastic dynamics of the model when designing the optimal controller may no longer be valid when the system is deployed in the real-world. The problem addressed by this paper is the following: Suppose we obtain an optimal trajectory by solving a control problem in the training environment, how do we ensure that the real-world system trajectory tracks this optimal trajectory with minimal amount of error in a deployment environment. In other words, we want to learn how we can adapt an optimal trained policy to distribution shifts in the environment. Distribution shifts are problematic in safety-critical systems, where a trained policy may lead to unsafe outcomes during deployment. We show that this problem can be cast as a nonlinear optimization problem that could be solved using heuristic method such as particle swarm optimization (PSO). However, if we instead consider a convex relaxation of this problem, we can learn policies that track the optimal trajectory with much better error performance, and faster computation times. We demonstrate the efficacy of our approach on tracking an optimal path using a Dubin's car model, and collision avoidance using both a linear and nonlinear model for adaptive cruise control.

翻译：许多真实世界的系统往往包含高度非线性和不确定动力学的物理组件或环境。可以使用许多不同的控制算法为这种系统设计最优控制器，假设实际系统的模型具有相当高的保真度。然而，在实际部署系统时，设计最优控制器时对模型的随机动态所作的假设可能不再有效。本文讨论的问题是：假设我们通过在训练环境中解决控制问题来获得最优轨迹，我们如何确保在部署环境中，真实系统轨迹以最小的误差跟踪这个最优轨迹。换句话说，我们想知道如何适应最优训练策略来适应环境的分布转移。分布转移在安全关键系统中会带来问题，因为训练策略可能会在部署过程中导致不安全的结果。我们展示了这个问题可以被建模为一个非线性优化问题，可以用启发式方法如粒子群优化（PSO）来解决。但是，如果我们考虑这个问题的凸松弛，我们可以学习到跟踪最优轨迹的策略，具有更好的误差性能和更快的计算时间。我们在使用Dubin's car 模型跟踪最优路径和使用自适应巡航控制进行线性和非线性模型碰撞避免的场景中展示了我们方法的功效。

0

相关内容

分布转移

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

专知会员服务

37+阅读 · 2022年7月12日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【ICML2021】学习权衡不完美的示范

专知会员服务

15+阅读 · 2021年9月23日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【伯克利-Ke Li】学习优化，74页ppt，Learning to Optimize

【伯克利-Ke Li】学习优化，74页ppt，Learning to Optimize

专知会员服务

41+阅读 · 2020年7月23日

【CVPR2020-Oral】无监督域内自适应语义分割，Unsupervised Intra-domain Adaptation

【CVPR2020-Oral】无监督域内自适应语义分割，Unsupervised Intra-domain Adaptation

专知会员服务

71+阅读 · 2020年4月20日

2019必读的十大深度强化学习论文

2019必读的十大深度强化学习论文

专知会员服务

59+阅读 · 2020年1月16日

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

专知会员服务

16+阅读 · 2019年12月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【ALT 2019 Tutorials】强化学习的探索性开发（Exploration-Exploitation in Reinforcement Learning）

【ALT 2019 Tutorials】强化学习的探索性开发（Exploration-Exploitation in Reinforcement Learning）

专知会员服务

34+阅读 · 2019年3月21日

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

专知

2+阅读 · 2022年7月12日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

利用动态深度学习预测金融时间序列基于Python

利用动态深度学习预测金融时间序列基于Python

量化投资与机器学习

18+阅读 · 2018年10月30日

【论文推荐】最新七篇强化学习相关论文—逻辑约束、综述、多任务深度强化学习、参数服务器、事件抽取、分层强化学习、过拟合研究

【论文推荐】最新七篇强化学习相关论文—逻辑约束、综述、多任务深度强化学习、参数服务器、事件抽取、分层强化学习、过拟合研究

专知

25+阅读 · 2018年4月29日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

I型干扰素抑制肿瘤转移的功能与分子机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

迭代变化因素下基于二维H∞理论的迭代学习控制方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

随机环境下卡尔曼滤波器动态特性

国家自然科学基金

1+阅读 · 2012年12月31日

多智能体不确定性系统的自适应一致性问题研究

国家自然科学基金

6+阅读 · 2012年12月31日

受限制策略下多臂Bandit过程的理论与应用研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于非因果稳定逆的柔性机械臂学习控制

国家自然科学基金

0+阅读 · 2012年12月31日

基于策略迭代算法的随机Markov跳变系统优化控制研究

国家自然科学基金

0+阅读 · 2012年12月31日

非局部模型的自适应算法研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于动态分层与自学习的多智能体自适应协作模型

国家自然科学基金

17+阅读 · 2008年12月31日

不确定环境下铁路集装箱动态多阶段调运优化模型和算法研究

国家自然科学基金

0+阅读 · 2008年12月31日

Adaptive Policy Learning to Additional Tasks

Arxiv

0+阅读 · 2023年5月24日

On Context Distribution Shift in Task Representation Learning for Offline Meta RL

Arxiv

0+阅读 · 2023年5月23日

Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice

Arxiv

0+阅读 · 2023年5月22日

Adaptive action supervision in reinforcement learning from real-world multi-agent demonstrations

Arxiv

0+阅读 · 2023年5月22日

Quantification before Selection: Active Dynamics Preference for Robust Reinforcement Learning

Arxiv

0+阅读 · 2023年5月20日

A General Framework for Fair Allocation under Matroid Rank Valuations

Arxiv

0+阅读 · 2023年5月20日

Distributional Multi-Objective Decision Making

Arxiv

0+阅读 · 2023年5月19日

Distributionally Robust Bayesian Optimization with $φ$-divergences

Arxiv

0+阅读 · 2023年5月19日

Introduction to Online Convex Optimization

Arxiv

23+阅读 · 2021年12月19日

Active Learning for Domain Adaptation: An Energy-based Approach

Arxiv

13+阅读 · 2021年12月2日

VIP会员

文章信息

相关主题

相关VIP内容

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

专知会员服务

37+阅读 · 2022年7月12日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【ICML2021】学习权衡不完美的示范

专知会员服务

15+阅读 · 2021年9月23日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【伯克利-Ke Li】学习优化，74页ppt，Learning to Optimize

【伯克利-Ke Li】学习优化，74页ppt，Learning to Optimize

专知会员服务

41+阅读 · 2020年7月23日

【CVPR2020-Oral】无监督域内自适应语义分割，Unsupervised Intra-domain Adaptation

【CVPR2020-Oral】无监督域内自适应语义分割，Unsupervised Intra-domain Adaptation

专知会员服务

71+阅读 · 2020年4月20日

2019必读的十大深度强化学习论文

2019必读的十大深度强化学习论文

专知会员服务

59+阅读 · 2020年1月16日

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

【Facebook|AAAI2020】在合作的部分可观察博弈中通过搜索改进策略（Improving Policies via Search in Cooperative Partially Observable Games）

专知会员服务

16+阅读 · 2019年12月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【ALT 2019 Tutorials】强化学习的探索性开发（Exploration-Exploitation in Reinforcement Learning）

【ALT 2019 Tutorials】强化学习的探索性开发（Exploration-Exploitation in Reinforcement Learning）

专知会员服务

34+阅读 · 2019年3月21日

热门VIP内容

开通专知VIP会员享更多权益服务

【伯克利博士论文】通过真实世界实践赋能机器人自主性

军用无人机集群技术尚未成熟——但潜力可期

人工智能安全治理白皮书（2025）

AgentOps综述：分类、挑战与未来方向

相关资讯

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

专知

2+阅读 · 2022年7月12日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

利用动态深度学习预测金融时间序列基于Python

利用动态深度学习预测金融时间序列基于Python

量化投资与机器学习

18+阅读 · 2018年10月30日

【论文推荐】最新七篇强化学习相关论文—逻辑约束、综述、多任务深度强化学习、参数服务器、事件抽取、分层强化学习、过拟合研究

【论文推荐】最新七篇强化学习相关论文—逻辑约束、综述、多任务深度强化学习、参数服务器、事件抽取、分层强化学习、过拟合研究

专知

25+阅读 · 2018年4月29日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

相关论文

Adaptive Policy Learning to Additional Tasks

Arxiv

0+阅读 · 2023年5月24日

On Context Distribution Shift in Task Representation Learning for Offline Meta RL

Arxiv

0+阅读 · 2023年5月23日

Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice

Arxiv

0+阅读 · 2023年5月22日

Adaptive action supervision in reinforcement learning from real-world multi-agent demonstrations

Arxiv

0+阅读 · 2023年5月22日

Quantification before Selection: Active Dynamics Preference for Robust Reinforcement Learning

Arxiv

0+阅读 · 2023年5月20日

A General Framework for Fair Allocation under Matroid Rank Valuations

Arxiv

0+阅读 · 2023年5月20日

Distributional Multi-Objective Decision Making

Arxiv

0+阅读 · 2023年5月19日

Distributionally Robust Bayesian Optimization with $φ$-divergences

Arxiv

0+阅读 · 2023年5月19日

Introduction to Online Convex Optimization

Arxiv

23+阅读 · 2021年12月19日

Active Learning for Domain Adaptation: An Energy-based Approach

Arxiv

13+阅读 · 2021年12月2日

相关基金

I型干扰素抑制肿瘤转移的功能与分子机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

迭代变化因素下基于二维H∞理论的迭代学习控制方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

随机环境下卡尔曼滤波器动态特性

国家自然科学基金

1+阅读 · 2012年12月31日

多智能体不确定性系统的自适应一致性问题研究

国家自然科学基金

6+阅读 · 2012年12月31日

受限制策略下多臂Bandit过程的理论与应用研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于非因果稳定逆的柔性机械臂学习控制

国家自然科学基金

0+阅读 · 2012年12月31日

基于策略迭代算法的随机Markov跳变系统优化控制研究

国家自然科学基金

0+阅读 · 2012年12月31日

非局部模型的自适应算法研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于动态分层与自学习的多智能体自适应协作模型

国家自然科学基金

17+阅读 · 2008年12月31日

不确定环境下铁路集装箱动态多阶段调运优化模型和算法研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员