VT-CLIP:用视觉制导文字加强视觉-语言模型 (VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts) - 专知论文

会员服务 ·

0

Attention · 语义鸿沟 · MoDELS · INFORMS · 相关系数 ·

2022 年 11 月 3 日

VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts

翻译：VT-CLIP:用视觉制导文字加强视觉-语言模型

Longtian Qiu,Renrui Zhang,Ziyu Guo,Ziyao Zeng,Yafeng Li,Guangnan Zhang

Contrastive Language-Image Pre-training (CLIP) has drawn increasing attention recently for its transferable visual representation learning. However, due to the semantic gap within datasets, CLIP's pre-trained image-text alignment becomes sub-optimal on downstream tasks, which severely harms its transferring performance. To better adapt the cross-modality embedding space, we propose to enhance CLIP via Visual-guided Texts, named VT-CLIP. Specifically, we guide textual features of different categories to adaptively explore informative regions on the image and aggregate visual features by attention mechanisms. In this way, the texts become visual-guided, namely, more semantically correlated with downstream images, which greatly benefits the category-wise matching process. In few-shot settings, we evaluate our VT-CLIP on 11 well-known classification datasets to demonstrate its effectiveness.

翻译：然而,由于数据集中的语义差异,CLIP经过事先训练的图像文本调整在下游任务上变得不尽如人意,严重危害其转移性能。为了更好地调整跨模式嵌入空间,我们提议通过视觉指导文字(名为VT-CLIP)加强跨模式嵌入空间。具体地说,我们指导不同类别的文字特征,以便通过关注机制适应性地探索关于图像和综合视觉特征的信息区域。通过这种方式,文本成为视觉指导,即与下游图像更具语义相关性,这极大地有利于类别匹配进程。在几个场景中,我们用11个众所周知的分类数据集来评估我们的VT-CLIP,以展示其有效性。

0

相关内容

Attention

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

专知会员服务

12+阅读 · 2022年3月19日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

因果图，Causal Graphs，52页ppt

因果图，Causal Graphs，52页ppt

专知会员服务

250+阅读 · 2020年4月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

专知会员服务

50+阅读 · 2020年2月26日

【NLP模型的跨语言/跨领域迁移】《Transferring NLP models across languages and domains》

【NLP模型的跨语言/跨领域迁移】《Transferring NLP models across languages and domains》

专知会员服务

43+阅读 · 2019年11月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

裂隙岩体双重介质非达西渗流与应力耦合流变效应研究

国家自然科学基金

0+阅读 · 2015年12月31日

迭代变化因素下基于二维H∞理论的迭代学习控制方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

含缓冲层高强钢焊接接头疲劳扩展模型及机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

Fstl1在胰岛素抵抗中的作用及其机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

开挖卸荷诱致围岩失稳机理的透明模型试验研究

国家自然科学基金

0+阅读 · 2012年12月31日

亚微米尺度下Beta钛合金单晶力学行为及变形机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

利用Salinomycin研究肿瘤细胞自噬的分子机制

国家自然科学基金

0+阅读 · 2011年12月31日

裂纹在沿晶氧化膜内形核的应力腐蚀新机理

国家自然科学基金

0+阅读 · 2011年12月31日

柱状节理各向异性岩石力学特性及模型研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于爆轰气体-应力耦合作用的节理岩体爆破机理和控制研究

国家自然科学基金

0+阅读 · 2009年12月31日

Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

Arxiv

0+阅读 · 2022年12月22日

Critic-Guided Decoding for Controlled Text Generation

Arxiv

0+阅读 · 2022年12月21日

Recent Advances in Scene Image Representation and Classification

Arxiv

0+阅读 · 2022年12月21日

From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models

Arxiv

0+阅读 · 2022年12月21日

DuAT: Dual-Aggregation Transformer Network for Medical Image Segmentation

Arxiv

0+阅读 · 2022年12月21日

CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

Arxiv

0+阅读 · 2022年12月20日

Fast and efficient speech enhancement with variational autoencoders

Arxiv

0+阅读 · 2022年11月2日

A Survey on Generative Diffusion Model

Arxiv

46+阅读 · 2022年9月6日

Cross-Modal Self-Attention Network for Referring Image Segmentation

Cross-Modal Self-Attention Network for Referring Image Segmentation

Arxiv

18+阅读 · 2019年4月9日

Text Generation from Knowledge Graphs with Graph Transformers

Arxiv

35+阅读 · 2019年4月4日

VIP会员

文章信息

相关主题

相关VIP内容

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

专知会员服务

12+阅读 · 2022年3月19日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

因果图，Causal Graphs，52页ppt

因果图，Causal Graphs，52页ppt

专知会员服务

250+阅读 · 2020年4月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

专知会员服务

50+阅读 · 2020年2月26日

【NLP模型的跨语言/跨领域迁移】《Transferring NLP models across languages and domains》

【NLP模型的跨语言/跨领域迁移】《Transferring NLP models across languages and domains》

专知会员服务

43+阅读 · 2019年11月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《战区安全决策课程体系》最新244页

《"无人机航母"原型平台》

任务规划与地形分析：现代复杂环境作战导航体系

《攻击场景描述形式化模型研究》

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

相关论文

Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

Arxiv

0+阅读 · 2022年12月22日

Critic-Guided Decoding for Controlled Text Generation

Arxiv

0+阅读 · 2022年12月21日

Recent Advances in Scene Image Representation and Classification

Arxiv

0+阅读 · 2022年12月21日

From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models

Arxiv

0+阅读 · 2022年12月21日

DuAT: Dual-Aggregation Transformer Network for Medical Image Segmentation

Arxiv

0+阅读 · 2022年12月21日

CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

Arxiv

0+阅读 · 2022年12月20日

Fast and efficient speech enhancement with variational autoencoders

Arxiv

0+阅读 · 2022年11月2日

A Survey on Generative Diffusion Model

Arxiv

46+阅读 · 2022年9月6日

Cross-Modal Self-Attention Network for Referring Image Segmentation

Cross-Modal Self-Attention Network for Referring Image Segmentation

Arxiv

18+阅读 · 2019年4月9日

Text Generation from Knowledge Graphs with Graph Transformers

Arxiv

35+阅读 · 2019年4月4日

相关基金

裂隙岩体双重介质非达西渗流与应力耦合流变效应研究

国家自然科学基金

0+阅读 · 2015年12月31日

迭代变化因素下基于二维H∞理论的迭代学习控制方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

含缓冲层高强钢焊接接头疲劳扩展模型及机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

Fstl1在胰岛素抵抗中的作用及其机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

开挖卸荷诱致围岩失稳机理的透明模型试验研究

国家自然科学基金

0+阅读 · 2012年12月31日

亚微米尺度下Beta钛合金单晶力学行为及变形机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

利用Salinomycin研究肿瘤细胞自噬的分子机制

国家自然科学基金

0+阅读 · 2011年12月31日

裂纹在沿晶氧化膜内形核的应力腐蚀新机理

国家自然科学基金

0+阅读 · 2011年12月31日

柱状节理各向异性岩石力学特性及模型研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于爆轰气体-应力耦合作用的节理岩体爆破机理和控制研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员