连续录相专栏 (Consistent Video Instance Segmentation with Inter-Frame Recurrent Attention) - 专知论文

会员服务 ·

0

Attention · 示例 · Extensibility · MoDELS · 端到端 ·

2022 年 6 月 14 日

Consistent Video Instance Segmentation with Inter-Frame Recurrent Attention

翻译：连续录相专栏

Quanzeng You,Jiang Wang,Peng Chu,Andre Abrantes,Zicheng Liu

from arxiv, 11 pages, 5 figures, 4 tables

Video instance segmentation aims at predicting object segmentation masks for each frame, as well as associating the instances across multiple frames. Recent end-to-end video instance segmentation methods are capable of performing object segmentation and instance association together in a direct parallel sequence decoding/prediction framework. Although these methods generally predict higher quality object segmentation masks, they can fail to associate instances in challenging cases because they do not explicitly model the temporal instance consistency for adjacent frames. We propose a consistent end-to-end video instance segmentation framework with Inter-Frame Recurrent Attention to model both the temporal instance consistency for adjacent frames and the global temporal context. Our extensive experiments demonstrate that the Inter-Frame Recurrent Attention significantly improves temporal instance consistency while maintaining the quality of the object segmentation masks. Our model achieves state-of-the-art accuracy on both YouTubeVIS-2019 (62.1\%) and YouTubeVIS-2021 (54.7\%) datasets. In addition, quantitative and qualitative results show that the proposed methods predict more temporally consistent instance segmentation masks.

翻译：视频实例截断法旨在预测每个框架的物体分离面罩,以及将多个框架的情况联系起来。最近的端到端视频实例截断法能够在一个直接平行的序列解码/定位框架内,同时进行物体分离和实例关联。虽然这些方法一般预测物体分离面罩的质量较高,但无法在具有挑战性的案件中将情况联系起来,因为它们没有明确模拟相邻框架的时间实例一致性。我们提议一个一致的端到端视频实例截断框架,与跨频频频频谱经常注意模拟相邻框架和全球时间背景下的时间实例一致性。我们的广泛实验表明,频谱经常注意在保持物体分离面罩质量的同时,大大提高了时间实例的一致性。我们的模型在YouTubeVIS-2019 (62.1 ⁇ ) 和YouTubeVIS-2021 (54.7 ⁇ ⁇ ) 数据集上都实现了最新水平的精确度。此外,定量和定性结果显示,拟议的方法预测了时间一致性更强的实例分割面罩。

0

相关内容

Attention

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【CVPR 2022】基于Tracklet查询和建议的高效视频实例分割，Efficient Video Instance Segmentation via Tracklet Query and Proposal

【CVPR 2022】基于Tracklet查询和建议的高效视频实例分割，Efficient Video Instance Segmentation via Tracklet Query and Proposal

专知会员服务

16+阅读 · 2022年3月3日

【CVPR 2022】使用多模态Transformer的端到端视频对象分割，End-to-End Referring Video Object Segmentation with Multimodal Transformer

【CVPR 2022】使用多模态Transformer的端到端视频对象分割，End-to-End Referring Video Object Segmentation with Multimodal Transformer

专知会员服务

28+阅读 · 2022年3月3日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

TorchSeg：基于pytorch的语义分割算法开源了

TorchSeg：基于pytorch的语义分割算法开源了

极市平台

20+阅读 · 2019年1月28日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

专知

27+阅读 · 2018年2月7日

气液混合流体喷射器内两相混合机理及其多因素耦合机制

国家自然科学基金

0+阅读 · 2014年12月31日

新疆天山北坡经济带PM_2.5时空分布与LUCC的关联性研究

国家自然科学基金

0+阅读 · 2014年12月31日

靶向抑制TCTP在急性髓系白血病治疗中的作用及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

通航环境耦合作用的船舶动力系统能效提升灰色建模研究

国家自然科学基金

0+阅读 · 2014年12月31日

mTOR功能性单倍体通过ERS-IRE1/α-JNK通路调控乳腺癌细胞药物敏感性的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

可压缩湍流粒子输运的拉格朗日（Lagrangian）研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于稳定性约束的高效多相流连续-离散耦合模拟

国家自然科学基金

0+阅读 · 2013年12月31日

基于刚度模型的机器人误差建模及标定方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

基于时序InSAR的北京地区地面沉降对地下水开采的响应机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

IKK与JNK信号通路促进炎性衰老的分子机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

MinVIS: A Minimal Video Instance Segmentation Framework without Video-based Training

Arxiv

0+阅读 · 2022年8月3日

Information Prebuilt Recurrent Reconstruction Network for Video Super-Resolution

Information Prebuilt Recurrent Reconstruction Network for Video Super-Resolution

Arxiv

0+阅读 · 2022年8月3日

Per-Clip Video Object Segmentation

Arxiv

0+阅读 · 2022年8月3日

Texture based Prototypical Network for Few-Shot Semantic Segmentation of Forest Cover: Generalizing for Different Geographical Regions

Arxiv

0+阅读 · 2022年8月2日

Motion-aware Memory Network for Fast Video Salient Object Detection

Arxiv

0+阅读 · 2022年8月1日

ATCA: an Arc Trajectory Based Model with Curvature Attention for Video Frame Interpolation

ATCA: an Arc Trajectory Based Model with Curvature Attention for Video Frame Interpolation

Arxiv

0+阅读 · 2022年8月1日

Temporal Relational Modeling with Self-Supervision for Action Segmentation

Arxiv

13+阅读 · 2020年12月14日

End-to-End Dense Video Captioning with Masked Transformer

Arxiv

14+阅读 · 2018年4月3日

Video Captioning via Hierarchical Reinforcement Learning

Arxiv

20+阅读 · 2018年3月29日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

VIP会员

文章信息

相关主题

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【CVPR 2022】基于Tracklet查询和建议的高效视频实例分割，Efficient Video Instance Segmentation via Tracklet Query and Proposal

【CVPR 2022】基于Tracklet查询和建议的高效视频实例分割，Efficient Video Instance Segmentation via Tracklet Query and Proposal

专知会员服务

16+阅读 · 2022年3月3日

【CVPR 2022】使用多模态Transformer的端到端视频对象分割，End-to-End Referring Video Object Segmentation with Multimodal Transformer

【CVPR 2022】使用多模态Transformer的端到端视频对象分割，End-to-End Referring Video Object Segmentation with Multimodal Transformer

专知会员服务

28+阅读 · 2022年3月3日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美国海军陆战队软件定义网络应用案例：分布式防火墙自动化系统》148页

《多体环境下定位导航授时（PNT）系统研究》228页

软件定义无线电（SDR）：商业与军事领域的技术、应用及未来趋势

《攻势防空作战中无人追击者/规避者最优轨迹研究（含动态交战区建模）》95页

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

TorchSeg：基于pytorch的语义分割算法开源了

TorchSeg：基于pytorch的语义分割算法开源了

极市平台

20+阅读 · 2019年1月28日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

专知

27+阅读 · 2018年2月7日

相关论文

MinVIS: A Minimal Video Instance Segmentation Framework without Video-based Training

Arxiv

0+阅读 · 2022年8月3日

Information Prebuilt Recurrent Reconstruction Network for Video Super-Resolution

Information Prebuilt Recurrent Reconstruction Network for Video Super-Resolution

Arxiv

0+阅读 · 2022年8月3日

Per-Clip Video Object Segmentation

Arxiv

0+阅读 · 2022年8月3日

Texture based Prototypical Network for Few-Shot Semantic Segmentation of Forest Cover: Generalizing for Different Geographical Regions

Arxiv

0+阅读 · 2022年8月2日

Motion-aware Memory Network for Fast Video Salient Object Detection

Arxiv

0+阅读 · 2022年8月1日

ATCA: an Arc Trajectory Based Model with Curvature Attention for Video Frame Interpolation

ATCA: an Arc Trajectory Based Model with Curvature Attention for Video Frame Interpolation

Arxiv

0+阅读 · 2022年8月1日

Temporal Relational Modeling with Self-Supervision for Action Segmentation

Arxiv

13+阅读 · 2020年12月14日

End-to-End Dense Video Captioning with Masked Transformer

Arxiv

14+阅读 · 2018年4月3日

Video Captioning via Hierarchical Reinforcement Learning

Arxiv

20+阅读 · 2018年3月29日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

相关基金

气液混合流体喷射器内两相混合机理及其多因素耦合机制

国家自然科学基金

0+阅读 · 2014年12月31日

新疆天山北坡经济带PM_2.5时空分布与LUCC的关联性研究

国家自然科学基金

0+阅读 · 2014年12月31日

靶向抑制TCTP在急性髓系白血病治疗中的作用及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

通航环境耦合作用的船舶动力系统能效提升灰色建模研究

国家自然科学基金

0+阅读 · 2014年12月31日

mTOR功能性单倍体通过ERS-IRE1/α-JNK通路调控乳腺癌细胞药物敏感性的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

可压缩湍流粒子输运的拉格朗日（Lagrangian）研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于稳定性约束的高效多相流连续-离散耦合模拟

国家自然科学基金

0+阅读 · 2013年12月31日

基于刚度模型的机器人误差建模及标定方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

基于时序InSAR的北京地区地面沉降对地下水开采的响应机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

IKK与JNK信号通路促进炎性衰老的分子机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员