重新审视线性二次调节控制：基于滚动视角策略梯度的视角 (Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient) - 专知论文

会员服务 ·

0

策略梯度 · 控制策略 · 梯度 · 稳定控制 · 卡尔曼滤波器 ·

2023 年 4 月 6 日

Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient

翻译：重新审视线性二次调节控制：基于滚动视角策略梯度的视角

Xiangyuan Zhang,Tamer Başar

We revisit in this paper the discrete-time linear quadratic regulator (LQR) problem from the perspective of receding-horizon policy gradient (RHPG), a newly developed model-free learning framework for control applications. We provide a fine-grained sample complexity analysis for RHPG to learn a control policy that is both stabilizing and $\epsilon$-close to the optimal LQR solution, and our algorithm does not require knowing a stabilizing control policy for initialization. Combined with the recent application of RHPG in learning the Kalman filter, we demonstrate the general applicability of RHPG in linear control and estimation with streamlined analyses.

翻译：本文从滚动视角策略梯度 (RHPG) 的角度重新审视了离散时间线性二次调节器(LQR)问题。我们为RHPG提供了精细的样本复杂度分析，以学习同时具有稳定性和$\epsilon $ -接近最优LQR解的控制策略，且我们的算法不需要知道初始化的稳定控制策略。结合RHPG在学习卡尔曼滤波器上的最新应用，我们展示了RHPG在线性控制和估计中的一般适用性，并提供了简化的分析方法。

0

相关内容

策略梯度

【CTH博士论文】基于强化学习的自动驾驶决策，149页pdf

【CTH博士论文】基于强化学习的自动驾驶决策，149页pdf

专知会员服务

58+阅读 · 2023年2月18日

【NeurIPS2021】非凸从动件的基于梯度的双层优化

专知会员服务

13+阅读 · 2021年10月12日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

专知会员服务

82+阅读 · 2020年2月27日

MIT-深度学习Deep Learning State of the Art in 2020，87页ppt

MIT-深度学习Deep Learning State of the Art in 2020，87页ppt

专知会员服务

62+阅读 · 2020年2月17日

实时强化学习《Real-Time Reinforcement Learning》S Ramstedt, C Pal [Mila, Element AI] (2019)

实时强化学习《Real-Time Reinforcement Learning》S Ramstedt, C Pal [Mila, Element AI] (2019)

专知会员服务

13+阅读 · 2019年11月17日

【CoRL2019最佳论文】模仿学习，A Divergence Minimization Perspective on Imitation Learning Methods

【CoRL2019最佳论文】模仿学习，A Divergence Minimization Perspective on Imitation Learning Methods

专知会员服务

24+阅读 · 2019年11月11日

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

专知会员服务

246+阅读 · 2019年10月21日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

OpenAI官方发布：强化学习中的关键论文

OpenAI官方发布：强化学习中的关键论文

专知

14+阅读 · 2018年12月12日

【OpenAI】深度强化学习关键论文列表

【OpenAI】深度强化学习关键论文列表

专知

11+阅读 · 2018年11月10日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

非线性ODE-PDE耦合系统的模糊建模与控制

国家自然科学基金

0+阅读 · 2014年12月31日

基于强化学习的前列腺癌蛋白质间相互作用网络的模型及方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

随机多智能体系统的一致性及优化控制

国家自然科学基金

1+阅读 · 2013年12月31日

时滞分数阶系统的控制器鲁棒稳定域及鲁棒控制

国家自然科学基金

0+阅读 · 2013年12月31日

基于粗糙信息的多自主体编队控制

国家自然科学基金

4+阅读 · 2013年12月31日

多智能体不确定性系统的自适应一致性问题研究

国家自然科学基金

6+阅读 · 2012年12月31日

促性腺激素调节C-型钠肽及其受体表达的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

回转旋臂式船用起重机系统动力学建模与非线性控制

国家自然科学基金

0+阅读 · 2012年12月31日

实时效率最优的感应电机无差拍直接转矩控制研究

国家自然科学基金

0+阅读 · 2009年12月31日

有界噪声激励下非线性系统的全局动力学研究

国家自然科学基金

0+阅读 · 2008年12月31日

State-Based $\infty$P-Set Conflict-Free Replicated Data Type

Arxiv

0+阅读 · 2023年5月26日

Option-Aware Adversarial Inverse Reinforcement Learning for Robotic Control

Arxiv

0+阅读 · 2023年5月26日

Physics-Guided Discovery of Highly Nonlinear Parametric Partial Differential Equations

Arxiv

0+阅读 · 2023年5月26日

An Analysis of Quantile Temporal-Difference Learning

Arxiv

0+阅读 · 2023年5月25日

Concurrent Constrained Optimization of Unknown Rewards for Multi-Robot Task Allocation

Arxiv

0+阅读 · 2023年5月24日

Policy Learning based on Deep Koopman Representation

Arxiv

0+阅读 · 2023年5月24日

Provable Offline Reinforcement Learning with Human Feedback

Arxiv

0+阅读 · 2023年5月24日

A Survey of Deep Reinforcement Learning in Recommender Systems: A Systematic Review and Future Directions

Arxiv

15+阅读 · 2021年9月8日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

A Survey of Deep Learning for Scientific Discovery

A Survey of Deep Learning for Scientific Discovery

Arxiv

29+阅读 · 2020年3月26日

VIP会员

文章信息

相关主题

卡尔曼滤波器

相关VIP内容

【CTH博士论文】基于强化学习的自动驾驶决策，149页pdf

【CTH博士论文】基于强化学习的自动驾驶决策，149页pdf

专知会员服务

58+阅读 · 2023年2月18日

【NeurIPS2021】非凸从动件的基于梯度的双层优化

专知会员服务

13+阅读 · 2021年10月12日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

专知会员服务

82+阅读 · 2020年2月27日

MIT-深度学习Deep Learning State of the Art in 2020，87页ppt

MIT-深度学习Deep Learning State of the Art in 2020，87页ppt

专知会员服务

62+阅读 · 2020年2月17日

实时强化学习《Real-Time Reinforcement Learning》S Ramstedt, C Pal [Mila, Element AI] (2019)

实时强化学习《Real-Time Reinforcement Learning》S Ramstedt, C Pal [Mila, Element AI] (2019)

专知会员服务

13+阅读 · 2019年11月17日

【CoRL2019最佳论文】模仿学习，A Divergence Minimization Perspective on Imitation Learning Methods

【CoRL2019最佳论文】模仿学习，A Divergence Minimization Perspective on Imitation Learning Methods

专知会员服务

24+阅读 · 2019年11月11日

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

专知会员服务

246+阅读 · 2019年10月21日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

热门VIP内容

开通专知VIP会员享更多权益服务

【博士论文】面向开放式世界的鲁棒智能体

美空军如何利用人工智能提升其兵棋推演能力

【AAAI2026】NeSTR：一种用于大型语言模型的神经-符号可溯因框架，用于时间推理

深度强化学习与模仿学习导论

相关资讯

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

OpenAI官方发布：强化学习中的关键论文

OpenAI官方发布：强化学习中的关键论文

专知

14+阅读 · 2018年12月12日

【OpenAI】深度强化学习关键论文列表

【OpenAI】深度强化学习关键论文列表

专知

11+阅读 · 2018年11月10日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

State-Based $\infty$P-Set Conflict-Free Replicated Data Type

Arxiv

0+阅读 · 2023年5月26日

Option-Aware Adversarial Inverse Reinforcement Learning for Robotic Control

Arxiv

0+阅读 · 2023年5月26日

Physics-Guided Discovery of Highly Nonlinear Parametric Partial Differential Equations

Arxiv

0+阅读 · 2023年5月26日

An Analysis of Quantile Temporal-Difference Learning

Arxiv

0+阅读 · 2023年5月25日

Concurrent Constrained Optimization of Unknown Rewards for Multi-Robot Task Allocation

Arxiv

0+阅读 · 2023年5月24日

Policy Learning based on Deep Koopman Representation

Arxiv

0+阅读 · 2023年5月24日

Provable Offline Reinforcement Learning with Human Feedback

Arxiv

0+阅读 · 2023年5月24日

A Survey of Deep Reinforcement Learning in Recommender Systems: A Systematic Review and Future Directions

Arxiv

15+阅读 · 2021年9月8日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

A Survey of Deep Learning for Scientific Discovery

A Survey of Deep Learning for Scientific Discovery

Arxiv

29+阅读 · 2020年3月26日

相关基金

非线性ODE-PDE耦合系统的模糊建模与控制

国家自然科学基金

0+阅读 · 2014年12月31日

基于强化学习的前列腺癌蛋白质间相互作用网络的模型及方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

随机多智能体系统的一致性及优化控制

国家自然科学基金

1+阅读 · 2013年12月31日

时滞分数阶系统的控制器鲁棒稳定域及鲁棒控制

国家自然科学基金

0+阅读 · 2013年12月31日

基于粗糙信息的多自主体编队控制

国家自然科学基金

4+阅读 · 2013年12月31日

多智能体不确定性系统的自适应一致性问题研究

国家自然科学基金

6+阅读 · 2012年12月31日

促性腺激素调节C-型钠肽及其受体表达的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

回转旋臂式船用起重机系统动力学建模与非线性控制

国家自然科学基金

0+阅读 · 2012年12月31日

实时效率最优的感应电机无差拍直接转矩控制研究

国家自然科学基金

0+阅读 · 2009年12月31日

有界噪声激励下非线性系统的全局动力学研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员