CLIP-PAE: 嵌入提取相关特征的预测-增强,以解析、可解释和可控文本制导图像操纵 (CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Image Manipulation) - 专知论文

会员服务 ·

0

控制器 · 相关特征 · 优化器 · Subspace · state-of-the-art ·

2022 年 11 月 25 日

CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Image Manipulation

翻译：CLIP-PAE: 嵌入提取相关特征的预测-增强,以解析、可解释和可控文本制导图像操纵

Chenliang Zhou,Fangcheng Zhong,Cengiz Oztireli

Recently introduced Contrastive Language-Image Pre-Training (CLIP) bridges images and text by embedding them into a joint latent space. This opens the door to ample literature that aims to manipulate an input image by providing a textual explanation. However, due to the discrepancy between image and text embeddings in the joint space, using text embeddings as the optimization target often introduces undesired artifacts in the resulting images. Disentanglement, interpretability, and controllability are also hard to guarantee for manipulation. To alleviate these problems, we propose to define corpus subspaces spanned by relevant prompts to capture specific image characteristics. We introduce CLIP Projection-Augmentation Embedding (PAE) as an optimization target to improve the performance of text-guided image manipulation. Our method is a simple and general paradigm that can be easily computed and adapted, and smoothly incorporated into any CLIP-based image manipulation algorithm. To demonstrate the effectiveness of our method, we conduct several theoretical and empirical studies. As a case study, we utilize the method for text-guided semantic face editing. We quantitatively and qualitatively demonstrate that PAE facilitates a more disentangled, interpretable, and controllable image manipulation with state-of-the-art quality and accuracy.

翻译：最近引入了不理想语言图像培训前(CLIP)的桥梁图像和文本,将其嵌入共同潜伏空间。这为大量文献打开了大门,这些文献旨在通过提供文本解释来操纵输入图像。然而,由于图像和文本嵌入于共同空间,使用文字嵌入作为优化目标,常常在生成图像中引入不受欢迎的文物。分解、可解释性和可控制性也难以保证操作。为了缓解这些问题,我们提议用相关提示来定义物质子空间,以获取特定图像特征。我们引入了CLIP预测增强嵌入(PAE)作为优化目标,以改善文本制导图像操纵的性能。我们的方法是一个简单和一般的范例,可以很容易地进行计算和调整,并顺利地融入基于CLIP的图像操纵算法。为了证明我们的方法的有效性,我们进行了一些理论和经验研究。作为案例研究,我们使用了文本制导图像面编辑的方法。我们引入了一种优化目标,我们用定量和定性的精确度来解释和定性地展示了PAEAE的准确性。我们用一种更难理解和定性的状态的图像。

0

相关内容

控制器

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【北京大学】探索提取跨模态信息进行图像caption，Exploring and Distilling Cross-Modal Information for Image Captioning

【北京大学】探索提取跨模态信息进行图像caption，Exploring and Distilling Cross-Modal Information for Image Captioning

专知会员服务

54+阅读 · 2020年3月3日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新六篇知识图谱相关论文—Zero-shot识别、卷积二维知识图谱、变分知识图谱推理、张量分解、推荐

【论文推荐】最新六篇知识图谱相关论文—Zero-shot识别、卷积二维知识图谱、变分知识图谱推理、张量分解、推荐

专知

50+阅读 · 2018年4月25日

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

专知

27+阅读 · 2018年2月7日

石墨烯/六方氮化硼面内拼接超晶格结构的可控制备及物性研究

国家自然科学基金

0+阅读 · 2015年12月31日

帕金森病酰胺质子转移磁共振成像研究

国家自然科学基金

0+阅读 · 2013年12月31日

非晶/纳米晶复合材料原子尺度塑性机制

国家自然科学基金

0+阅读 · 2013年12月31日

还原敏感触发式纳米Pickering乳递药系统的构建及靶向肝癌的研究

国家自然科学基金

0+阅读 · 2012年12月31日

DNA介导钌配合物包覆纳米颗粒的新型光电复合材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于多三氮大环过渡金属配合物的顺磁CEST磁共振成像对比剂的制备及其对比效能研究

国家自然科学基金

0+阅读 · 2012年12月31日

微小RNA-375降低心肌成纤维细胞IL-33表达在糖尿病心肌病发病中的作用

国家自然科学基金

0+阅读 · 2011年12月31日

强电场作用下陶瓷/金属扩散连接界面反应及机理

国家自然科学基金

0+阅读 · 2011年12月31日

新型方酸衍生物类抗肿瘤药物IMB-13的结构改造和作用机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

热休克转录因子 1 作为癌靶蛋白在肝细胞癌中的分子调控机理

国家自然科学基金

0+阅读 · 2009年12月31日

A Multi-View Joint Learning Framework for Embedding Clinical Codes and Text Using Graph Neural Networks

Arxiv

0+阅读 · 2023年1月27日

Mixed Attention Network for Hyperspectral Image Denoising

Arxiv

0+阅读 · 2023年1月27日

Optimizing Feature Set for Click-Through Rate Prediction

Arxiv

0+阅读 · 2023年1月26日

Survey: Image Mixing and Deleting for Data Augmentation

Arxiv

0+阅读 · 2023年1月25日

Diverse Single Image Generation with Controllable Global Structure

Arxiv

0+阅读 · 2023年1月25日

In Which Graph Structures Can We Efficiently Find Temporally Disjoint Paths and Walks?

Arxiv

0+阅读 · 2023年1月25日

From Show to Tell: A Survey on Image Captioning

Arxiv

15+阅读 · 2021年7月14日

Graph Enhanced Representation Learning for News Recommendation

Arxiv

24+阅读 · 2020年3月31日

On Feature Normalization and Data Augmentation

On Feature Normalization and Data Augmentation

Arxiv

15+阅读 · 2020年2月25日

Zero-shot Recognition via Semantic Embeddings and Knowledge Graphs

Arxiv

18+阅读 · 2018年4月8日

VIP会员

文章信息

相关主题

state-of-the-art

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【北京大学】探索提取跨模态信息进行图像caption，Exploring and Distilling Cross-Modal Information for Image Captioning

【北京大学】探索提取跨模态信息进行图像caption，Exploring and Distilling Cross-Modal Information for Image Captioning

专知会员服务

54+阅读 · 2020年3月3日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【新书】面向企业的图学习扩展：生产级图学习与推理，485页pdf

AI智能体编程：技术、挑战与机遇综述

【国家标准】数据安全技术数据安全风险评估方法

【CMU博士论文】交互式学习的进展：替代性反馈机制与自适应因果推理

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新六篇知识图谱相关论文—Zero-shot识别、卷积二维知识图谱、变分知识图谱推理、张量分解、推荐

【论文推荐】最新六篇知识图谱相关论文—Zero-shot识别、卷积二维知识图谱、变分知识图谱推理、张量分解、推荐

专知

50+阅读 · 2018年4月25日

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

专知

27+阅读 · 2018年2月7日

相关论文

A Multi-View Joint Learning Framework for Embedding Clinical Codes and Text Using Graph Neural Networks

Arxiv

0+阅读 · 2023年1月27日

Mixed Attention Network for Hyperspectral Image Denoising

Arxiv

0+阅读 · 2023年1月27日

Optimizing Feature Set for Click-Through Rate Prediction

Arxiv

0+阅读 · 2023年1月26日

Survey: Image Mixing and Deleting for Data Augmentation

Arxiv

0+阅读 · 2023年1月25日

Diverse Single Image Generation with Controllable Global Structure

Arxiv

0+阅读 · 2023年1月25日

In Which Graph Structures Can We Efficiently Find Temporally Disjoint Paths and Walks?

Arxiv

0+阅读 · 2023年1月25日

From Show to Tell: A Survey on Image Captioning

Arxiv

15+阅读 · 2021年7月14日

Graph Enhanced Representation Learning for News Recommendation

Arxiv

24+阅读 · 2020年3月31日

On Feature Normalization and Data Augmentation

On Feature Normalization and Data Augmentation

Arxiv

15+阅读 · 2020年2月25日

Zero-shot Recognition via Semantic Embeddings and Knowledge Graphs

Arxiv

18+阅读 · 2018年4月8日

相关基金

石墨烯/六方氮化硼面内拼接超晶格结构的可控制备及物性研究

国家自然科学基金

0+阅读 · 2015年12月31日

帕金森病酰胺质子转移磁共振成像研究

国家自然科学基金

0+阅读 · 2013年12月31日

非晶/纳米晶复合材料原子尺度塑性机制

国家自然科学基金

0+阅读 · 2013年12月31日

还原敏感触发式纳米Pickering乳递药系统的构建及靶向肝癌的研究

国家自然科学基金

0+阅读 · 2012年12月31日

DNA介导钌配合物包覆纳米颗粒的新型光电复合材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于多三氮大环过渡金属配合物的顺磁CEST磁共振成像对比剂的制备及其对比效能研究

国家自然科学基金

0+阅读 · 2012年12月31日

微小RNA-375降低心肌成纤维细胞IL-33表达在糖尿病心肌病发病中的作用

国家自然科学基金

0+阅读 · 2011年12月31日

强电场作用下陶瓷/金属扩散连接界面反应及机理

国家自然科学基金

0+阅读 · 2011年12月31日

新型方酸衍生物类抗肿瘤药物IMB-13的结构改造和作用机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

热休克转录因子 1 作为癌靶蛋白在肝细胞癌中的分子调控机理

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员