使用预先培训的愿景语言模型的开放式词汇语义分解简单基线 (A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model) - 专知论文

会员服务 ·

0

Performer · SimPLe · MoDELS · FCN · 基准 ·

2022 年 12 月 29 日

A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model

翻译：使用预先培训的愿景语言模型的开放式词汇语义分解简单基线

Mengde Xu,Zheng Zhang,Fangyun Wei,Yutong Lin,Yue Cao,Han Hu,Xiang Bai

Recently, open-vocabulary image classification by vision language pre-training has demonstrated incredible achievements, that the model can classify arbitrary categories without seeing additional annotated images of that category. However, it is still unclear how to make the open-vocabulary recognition work well on broader vision problems. This paper targets open-vocabulary semantic segmentation by building it on an off-the-shelf pre-trained vision-language model, i.e., CLIP. However, semantic segmentation and the CLIP model perform on different visual granularity, that semantic segmentation processes on pixels while CLIP performs on images. To remedy the discrepancy in processing granularity, we refuse the use of the prevalent one-stage FCN based framework, and advocate a two-stage semantic segmentation framework, with the first stage extracting generalizable mask proposals and the second stage leveraging an image based CLIP model to perform open-vocabulary classification on the masked image crops which are generated in the first stage. Our experimental results show that this two-stage framework can achieve superior performance than FCN when trained only on COCO Stuff dataset and evaluated on other datasets without fine-tuning. Moreover, this simple framework also surpasses previous state-of-the-arts of zero-shot semantic segmentation by a large margin: +29.5 hIoU on the Pascal VOC 2012 dataset, and +8.9 hIoU on the COCO Stuff dataset. With its simplicity and strong performance, we hope this framework to serve as a baseline to facilitate future research. The code are made publicly available at~\url{https://github.com/MendelXu/zsseg.baseline}.

翻译：最近,通过视觉语言培训前的开放式语言图像分类 3 展示了令人难以置信的成就,该模型可以在不看到该类别附加附加说明的图像的情况下对任意分类进行分类,然而,目前还不清楚如何使开放式语言识别在更广泛的视觉问题上发挥良好的作用。本文的目标是开放语言语义的语义分解, 将其建在现成的预先培训的视觉语言模型上, 即 CLIP。然而, 语义分解和 CLIP 模型在不同视觉颗粒上运行, 在像素上的语义分解过程, 而 CLIP 则在图像上运行。然而, 为了纠正处理颗粒中的差异, 我们拒绝使用流行的一阶段FCN 框架, 并倡导两阶段语义的语义分解框架, 其第一阶段是提取通用的面面面面语义建议, 第二阶段是利用基于 CLIP 模型对面面面语系图像进行公开分类。我们的实验结果显示, 这个二阶段框架可以实现更高的性性能, 而不是FCN 。

0

相关内容

Performer

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

【ICML2020投稿论文】用于半监督图像分类的CowMask，Milking CowMask for Semi-Supervised Image Classification

【ICML2020投稿论文】用于半监督图像分类的CowMask，Milking CowMask for Semi-Supervised Image Classification

专知会员服务

29+阅读 · 2020年3月27日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

一类离散Hindmarsh-Rose模型的分支延拓

国家自然科学基金

0+阅读 · 2015年12月31日

芯-壳结构超高温多层陶瓷电容器介质材料的制备与性能调控

国家自然科学基金

0+阅读 · 2015年12月31日

复合级联多电平电池储能功率转换系统研究

国家自然科学基金

0+阅读 · 2013年12月31日

两类投资组合优化问题的模型与算法研究

国家自然科学基金

2+阅读 · 2013年12月31日

多尺度光伏玻璃压延成型机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于绝缘结构介电泳微流控器件中的电热流动

国家自然科学基金

0+阅读 · 2012年12月31日

多铁复合材料磁电耦合增强及界面结构调控研究

国家自然科学基金

0+阅读 · 2012年12月31日

多铁性LSCMO/PMN-PT磁电复合薄膜的制备、表征及原型器件探索

国家自然科学基金

0+阅读 · 2012年12月31日

钙钛矿型多铁性异质结的界面调控磁电耦合效应研究

国家自然科学基金

0+阅读 · 2011年12月31日

多铁性超晶格薄膜中电极化-电磁性-应力耦合的机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning

Arxiv

0+阅读 · 2023年2月28日

Internet Explorer: Targeted Representation Learning on the Open Web

Arxiv

0+阅读 · 2023年2月27日

Aligning Bag of Regions for Open-Vocabulary Object Detection

Arxiv

0+阅读 · 2023年2月27日

Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Arxiv

0+阅读 · 2023年2月24日

STA: Self-controlled Text Augmentation for Improving Text Classifications

Arxiv

0+阅读 · 2023年2月24日

Language-Driven Representation Learning for Robotics

Arxiv

1+阅读 · 2023年2月24日

F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Arxiv

0+阅读 · 2023年2月23日

Conditional Prompt Learning for Vision-Language Models

Conditional Prompt Learning for Vision-Language Models

Arxiv

13+阅读 · 2022年3月10日

Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph Completion

Arxiv

12+阅读 · 2021年4月27日

Making Pre-trained Language Models Better Few-shot Learners

Arxiv

14+阅读 · 2020年12月31日

VIP会员

文章信息

相关主题

相关VIP内容

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

【ICML2020投稿论文】用于半监督图像分类的CowMask，Milking CowMask for Semi-Supervised Image Classification

【ICML2020投稿论文】用于半监督图像分类的CowMask，Milking CowMask for Semi-Supervised Image Classification

专知会员服务

29+阅读 · 2020年3月27日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津博士论文】零样本强化学习综述

《美军条令：陆军指挥官与规划人员地理空间指南》60页

战术边缘指挥控制：防务面临的核心挑战

迈向开放世界检测：综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning

Arxiv

0+阅读 · 2023年2月28日

Internet Explorer: Targeted Representation Learning on the Open Web

Arxiv

0+阅读 · 2023年2月27日

Aligning Bag of Regions for Open-Vocabulary Object Detection

Arxiv

0+阅读 · 2023年2月27日

Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Arxiv

0+阅读 · 2023年2月24日

STA: Self-controlled Text Augmentation for Improving Text Classifications

Arxiv

0+阅读 · 2023年2月24日

Language-Driven Representation Learning for Robotics

Arxiv

1+阅读 · 2023年2月24日

F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Arxiv

0+阅读 · 2023年2月23日

Conditional Prompt Learning for Vision-Language Models

Conditional Prompt Learning for Vision-Language Models

Arxiv

13+阅读 · 2022年3月10日

Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph Completion

Arxiv

12+阅读 · 2021年4月27日

Making Pre-trained Language Models Better Few-shot Learners

Arxiv

14+阅读 · 2020年12月31日

相关基金

一类离散Hindmarsh-Rose模型的分支延拓

国家自然科学基金

0+阅读 · 2015年12月31日

芯-壳结构超高温多层陶瓷电容器介质材料的制备与性能调控

国家自然科学基金

0+阅读 · 2015年12月31日

复合级联多电平电池储能功率转换系统研究

国家自然科学基金

0+阅读 · 2013年12月31日

两类投资组合优化问题的模型与算法研究

国家自然科学基金

2+阅读 · 2013年12月31日

多尺度光伏玻璃压延成型机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于绝缘结构介电泳微流控器件中的电热流动

国家自然科学基金

0+阅读 · 2012年12月31日

多铁复合材料磁电耦合增强及界面结构调控研究

国家自然科学基金

0+阅读 · 2012年12月31日

多铁性LSCMO/PMN-PT磁电复合薄膜的制备、表征及原型器件探索

国家自然科学基金

0+阅读 · 2012年12月31日

钙钛矿型多铁性异质结的界面调控磁电耦合效应研究

国家自然科学基金

0+阅读 · 2011年12月31日

多铁性超晶格薄膜中电极化-电磁性-应力耦合的机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员