从次优示范中学习团队政策半模拟模拟学习 (Semi-Supervised Imitation Learning of Team Policies from Suboptimal Demonstrations) - 专知论文

会员服务 ·

0

TEAM · 学成 · Performer · 讲稿 · MoDELS ·

2022 年 5 月 5 日

Semi-Supervised Imitation Learning of Team Policies from Suboptimal Demonstrations

翻译：从次优示范中学习团队政策半模拟模拟学习

Sangwon Seo,Vaibhav V. Unhelkar

from arxiv, Extended version of an identically-titled paper accepted at IJCAI 2022

We present Bayesian Team Imitation Learner (BTIL), an imitation learning algorithm to model behavior of teams performing sequential tasks in Markovian domains. In contrast to existing multi-agent imitation learning techniques, BTIL explicitly models and infers the time-varying mental states of team members, thereby enabling learning of decentralized team policies from demonstrations of suboptimal teamwork. Further, to allow for sample- and label-efficient policy learning from small datasets, BTIL employs a Bayesian perspective and is capable of learning from semi-supervised demonstrations. We demonstrate and benchmark the performance of BTIL on synthetic multi-agent tasks as well as a novel dataset of human-agent teamwork. Our experiments show that BTIL can successfully learn team policies from demonstrations despite the influence of team members' (time-varying and potentially misaligned) mental states on their behavior.

翻译：我们介绍贝叶斯团队模拟学习者(BTIL),这是一种模拟学习算法,用以模拟在马尔科维亚地区执行连续任务的团队的行为。与现有的多试剂模拟学习技术相比,BTIL明确模型并推断了团队成员具有时间变化的心理状态,从而能够从次优团队协作的示范中学习分散的团队政策。此外,为了从小型数据集中学习样本和标签效率高的政策,BTIL采用了巴伊西亚视角,能够从半监督的演示中学习。我们展示并衡量BTIL在合成多试剂任务方面的表现以及人类代理团队合作的新数据集。我们的实验表明,尽管团队成员(时间变化和可能错配)精神状态对其行为产生了影响,但BTIL仍然能够成功地从演示中学习团队政策。

0

相关内容

TEAM

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

miR-124靶向TRAF6在骨肉瘤中的作用

国家自然科学基金

0+阅读 · 2013年12月31日

基于抑制Ti基体氧化的不连续薄膜形稳阳极

国家自然科学基金

0+阅读 · 2013年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

HIC1调控CIITA转录机制研究及其在B细胞分化中的意义

国家自然科学基金

0+阅读 · 2012年12月31日

基于应急预案管理的应急决策动态优化与协同机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

Stackelberg Risk Preference Design

Arxiv

0+阅读 · 2022年6月26日

Feedback Dynamics of the Low-Income Rental Housing Market: Exploring Policy Responses to COVID-19

Arxiv

0+阅读 · 2022年6月25日

Inference on the Best Policies with Many Covariates

Inference on the Best Policies with Many Covariates

Arxiv

0+阅读 · 2022年6月23日

Learning Agile Skills via Adversarial Imitation of Rough Partial Demonstrations

Arxiv

0+阅读 · 2022年6月23日

Latent Policies for Adversarial Imitation Learning

Arxiv

0+阅读 · 2022年6月22日

VIP会员

文章信息

相关主题

相关VIP内容

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《运用增强现实技术进行军事任务规划》130页

《高压决策环境中的人机协作》200页博士论文

《2025财年美陆军转型倡议（ATI）部队结构与组织提案》

《探索用于低层级任务区分与分类的转址旁路缓冲》

相关资讯

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

相关论文

Stackelberg Risk Preference Design

Arxiv

0+阅读 · 2022年6月26日

Feedback Dynamics of the Low-Income Rental Housing Market: Exploring Policy Responses to COVID-19

Arxiv

0+阅读 · 2022年6月25日

Inference on the Best Policies with Many Covariates

Inference on the Best Policies with Many Covariates

Arxiv

0+阅读 · 2022年6月23日

Learning Agile Skills via Adversarial Imitation of Rough Partial Demonstrations

Arxiv

0+阅读 · 2022年6月23日

Latent Policies for Adversarial Imitation Learning

Arxiv

0+阅读 · 2022年6月22日

相关基金

miR-124靶向TRAF6在骨肉瘤中的作用

国家自然科学基金

0+阅读 · 2013年12月31日

基于抑制Ti基体氧化的不连续薄膜形稳阳极

国家自然科学基金

0+阅读 · 2013年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

HIC1调控CIITA转录机制研究及其在B细胞分化中的意义

国家自然科学基金

0+阅读 · 2012年12月31日

基于应急预案管理的应急决策动态优化与协同机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员