南极顺序蒙特卡洛 (Critic Sequential Monte Carlo) - 专知论文

会员服务 ·

0

蒙特卡罗 · 评论员 · 回合 · 泛函 · 推断 ·

2023 年 1 月 21 日

Critic Sequential Monte Carlo

翻译：南极顺序蒙特卡洛

Vasileios Lioutas,Jonathan Wilder Lavington,Justice Sefas,Matthew Niedoba,Yunpeng Liu,Berend Zwartsenberg,Setareh Dabiri,Frank Wood,Adam Scibior

from arxiv, ICLR 2023

We introduce CriticSMC, a new algorithm for planning as inference built from a composition of sequential Monte Carlo with learned Soft-Q function heuristic factors. These heuristic factors, obtained from parametric approximations of the marginal likelihood ahead, more effectively guide SMC towards the desired target distribution, which is particularly helpful for planning in environments with hard constraints placed sparsely in time. Compared with previous work, we modify the placement of such heuristic factors, which allows us to cheaply propose and evaluate large numbers of putative action particles, greatly increasing inference and planning efficiency. CriticSMC is compatible with informative priors, whose density function need not be known, and can be used as a model-free control algorithm. Our experiments on collision avoidance in a high-dimensional simulated driving task show that CriticSMC significantly reduces collision rates at a low computational cost while maintaining realism and diversity of driving behaviors across vehicles and environment scenarios.

翻译：我们引入了CriticSMC(CriticSMC),这是一个规划的新算法,它由相继的蒙特卡洛构成,具有丰富的Soft-Q功能超常因素。这些超常因素来自未来边缘可能性的参数近似值,可以更有效地引导SMC实现预期的目标分布,这对于在困难环境中规划工作特别有帮助,而这种环境在时间上受到很少的制约。与以往的工作相比,我们修改了这种超常因素的位置,使我们能够廉价地提议和评价大量模拟动作粒子,大大提高了推断和规划效率。CriticSMC(CriticSMC)与信息前科相容,其密度功能不需要知道,可以用作无模型的控制算法。我们在高维模拟驾驶任务中避免碰撞的实验表明,CriticSMC(CriticSMC)在保持车辆和环境情景之间驾驶行为的现实主义和多样性的同时,以低计算成本大幅降低碰撞率。

0

相关内容

蒙特卡罗

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

暗共振的相干微扰与Bogoliubov变换研究

国家自然科学基金

0+阅读 · 2014年12月31日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

近似最优径向基函数插值的理论与算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

双溶剂高分子溶液的monte carlo研究

国家自然科学基金

0+阅读 · 2012年12月31日

时域连续的高维Monte Carlo绘制技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于点阵材料微观临界应力的几何多尺度拓扑优化方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

ICF中高能电子和离子输运的Monte-Carlo算法研究和程序研制

国家自然科学基金

0+阅读 · 2009年12月31日

土壤氮素行为及其模拟模型不确定性的Monte-Carlo分析

国家自然科学基金

0+阅读 · 2008年12月31日

表面等离子体共振单核苷酸多态性分型新方法研究

国家自然科学基金

0+阅读 · 2008年12月31日

几何动力学在非完整系统几何数值积分中的应用研究

国家自然科学基金

0+阅读 · 2008年12月31日

Quasi continuous level Monte Carlo for random elliptic PDEs

Arxiv

0+阅读 · 2023年3月15日

Replay Buffer With Local Forgetting for Adaptive Deep Model-Based Reinforcement Learning

Arxiv

0+阅读 · 2023年3月15日

Learning Minimally-Violating Continuous Control for Infeasible Linear Temporal Logic Specifications

Arxiv

0+阅读 · 2023年3月15日

A 2-opt Algorithm for Locally Optimal Set Partition Optimization

Arxiv

0+阅读 · 2023年3月14日

Learning Vehicle Trajectory Uncertainty

Arxiv

0+阅读 · 2023年3月13日

Low Frequency Spinning LiDAR De-Skewing

Low Frequency Spinning LiDAR De-Skewing

Arxiv

0+阅读 · 2023年3月13日

Real-time scheduling of renewable power systems through planning-based reinforcement learning

Arxiv

0+阅读 · 2023年3月13日

Analyzing Infrastructure LiDAR Placement with Realistic LiDAR Simulation Library

Arxiv

0+阅读 · 2023年3月11日

Doubly Robust Stein-Kernelized Monte Carlo Estimator: Simultaneous Bias-Variance Reduction and Supercanonical Convergence

Arxiv

0+阅读 · 2023年3月10日

Deep Reinforcement Learning for List-wise Recommendations

Arxiv

13+阅读 · 2018年1月5日

VIP会员

文章信息

相关主题

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《陆军战斗操练中的关键事件诊断》

《自适应训练辅助概念及其在空战管理员加速训练中的应用导论》最新126页

军事通信市场七大趋势概述

《抗干扰无人机蜂群行为的遗传算法方法》

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Quasi continuous level Monte Carlo for random elliptic PDEs

Arxiv

0+阅读 · 2023年3月15日

Replay Buffer With Local Forgetting for Adaptive Deep Model-Based Reinforcement Learning

Arxiv

0+阅读 · 2023年3月15日

Learning Minimally-Violating Continuous Control for Infeasible Linear Temporal Logic Specifications

Arxiv

0+阅读 · 2023年3月15日

A 2-opt Algorithm for Locally Optimal Set Partition Optimization

Arxiv

0+阅读 · 2023年3月14日

Learning Vehicle Trajectory Uncertainty

Arxiv

0+阅读 · 2023年3月13日

Low Frequency Spinning LiDAR De-Skewing

Low Frequency Spinning LiDAR De-Skewing

Arxiv

0+阅读 · 2023年3月13日

Real-time scheduling of renewable power systems through planning-based reinforcement learning

Arxiv

0+阅读 · 2023年3月13日

Analyzing Infrastructure LiDAR Placement with Realistic LiDAR Simulation Library

Arxiv

0+阅读 · 2023年3月11日

Doubly Robust Stein-Kernelized Monte Carlo Estimator: Simultaneous Bias-Variance Reduction and Supercanonical Convergence

Arxiv

0+阅读 · 2023年3月10日

Deep Reinforcement Learning for List-wise Recommendations

Arxiv

13+阅读 · 2018年1月5日

相关基金

暗共振的相干微扰与Bogoliubov变换研究

国家自然科学基金

0+阅读 · 2014年12月31日

Poisson流形上的修正Hamilton方法

国家自然科学基金

0+阅读 · 2014年12月31日

近似最优径向基函数插值的理论与算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

双溶剂高分子溶液的monte carlo研究

国家自然科学基金

0+阅读 · 2012年12月31日

时域连续的高维Monte Carlo绘制技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于点阵材料微观临界应力的几何多尺度拓扑优化方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

ICF中高能电子和离子输运的Monte-Carlo算法研究和程序研制

国家自然科学基金

0+阅读 · 2009年12月31日

土壤氮素行为及其模拟模型不确定性的Monte-Carlo分析

国家自然科学基金

0+阅读 · 2008年12月31日

表面等离子体共振单核苷酸多态性分型新方法研究

国家自然科学基金

0+阅读 · 2008年12月31日

几何动力学在非完整系统几何数值积分中的应用研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员