图片编辑器和EditBench：推进和评估文本引导的图像修复 (Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting) - 专知论文

会员服务 ·

0

图像修复 · 级联 · 属性 · 图像对齐 · 定量评估 ·

2023 年 4 月 12 日

Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting

翻译：图片编辑器和EditBench：推进和评估文本引导的图像修复

Su Wang,Chitwan Saharia,Ceslee Montgomery,Jordi Pont-Tuset,Shai Noy,Stefano Pellegrini,Yasumasa Onoe,Sarah Laszlo,David J. Fleet,Radu Soricut,Jason Baldridge,Mohammad Norouzi,Peter Anderson,William Chan

from arxiv, CVPR 2023 Camera Ready

Text-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen on text-guided image inpainting. Imagen Editor's edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment -- such that Imagen Editor is preferred over DALL-E 2 and Stable Diffusion -- and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes.

翻译：文本引导的图像编辑可以在支持创造性应用方面产生革命性的影响。其中一个关键挑战是生成忠实于输入文本提示的编辑，并与输入图像一致。我们提出了Imagen Editor，这是一个级联扩散模型，通过在文本引导的图像修复上对Imagen进行微调。Imagen Editor的编辑对文本提示忠实，这是通过在训练期间使用对象检测器提出修复蒙版来实现的。此外，Imagen Editor通过将级联流程条件化于原始高分辨率图像来捕获输入图像的细节。为了改善定性和定量评估，我们引入了EditBench，一个用于文本引导的图像修复的系统化评估基准。EditBench评估自然和生成图像上的修复，探索物体、属性和场景。通过对EditBench进行广泛的人类评估，我们发现在训练过程中使用对象掩模可导致文本-图像对齐的整体改进，使得Imagen Editor优于DALL-E 2和稳定扩散，并且这些模型作为同伴相对于文本渲染更擅长物体渲染，并且可以处理材料/颜色/大小属性而不是计数/形状属性。

0

相关内容

图像修复

图像修复（英语：Inpainting）指重建的图像和视频中丢失或损坏的部分的过程。例如在博物馆中，这项工作常由经验丰富的博物馆管理员或者艺术品修复师来进行。数码世界中，图像修复又称图像插值或视频插值，指利用复杂的算法来替换已丢失、损坏的图像数据，主要替换一些小区域和瑕疵。

【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型

【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型

专知会员服务

22+阅读 · 2022年12月5日

中科院自动化所17篇CVPR 2022 新作速览！

中科院自动化所17篇CVPR 2022 新作速览！

专知会员服务

20+阅读 · 2022年3月19日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【百度&北京大学】自然语言生成的保真性:分析、评价和优化方法的系统综述，Faithfulness in Natural Language Generation: A Systematic Survey of Analysis, Evaluation and Optimization Methods

【百度&北京大学】自然语言生成的保真性:分析、评价和优化方法的系统综述，Faithfulness in Natural Language Generation: A Systematic Survey of Analysis, Evaluation and Optimization Methods

专知会员服务

15+阅读 · 2022年3月11日

【南洋理工大学Chuanxia Zheng博士论文】基于深度生成学习的逼真图像合成，197页pdf，Synthesizing Photorealistic Images with Deep Generative Learning

【南洋理工大学Chuanxia Zheng博士论文】基于深度生成学习的逼真图像合成，197页pdf，Synthesizing Photorealistic Images with Deep Generative Learning

专知会员服务

20+阅读 · 2022年3月9日

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

专知会员服务

23+阅读 · 2022年3月3日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知会员服务

77+阅读 · 2020年7月23日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

专知会员服务

33+阅读 · 2020年2月29日

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

专知会员服务

43+阅读 · 2020年1月28日

Stable Diffusion再迎重磅更新！2.0版「涩图」功能被砍，网友狂打差评

Stable Diffusion再迎重磅更新！2.0版「涩图」功能被砍，网友狂打差评

新智元

0+阅读 · 2022年11月25日

7 Papers & Radios | 谷歌推出DreamBooth扩散模型；张益唐零点猜想论文出炉

7 Papers & Radios | 谷歌推出DreamBooth扩散模型；张益唐零点猜想论文出炉

机器之心

2+阅读 · 2022年11月13日

谷歌新作Imagen：用Transformer和扩散模型把"文字到图像生成"卷上天！

谷歌新作Imagen：用Transformer和扩散模型把"文字到图像生成"卷上天！

CVer

0+阅读 · 2022年5月27日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知

15+阅读 · 2020年7月23日

【清华出品】NLP新方向文本对抗攻击与防御必读论文列表

【清华出品】NLP新方向文本对抗攻击与防御必读论文列表

专知

21+阅读 · 2019年7月11日

已删除

将门创投

12+阅读 · 2019年7月1日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

实战 | 用Python做图像处理（一）

实战 | 用Python做图像处理（一）

七月在线实验室

25+阅读 · 2018年5月23日

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

专知

15+阅读 · 2018年5月15日

Generative Adversarial Text to Image Synthesis论文解读

Generative Adversarial Text to Image Synthesis论文解读

统计学习与视觉计算组

13+阅读 · 2017年6月9日

血管内窥光声图像的频率谱研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于多任务概率视觉语义模型的图像场景理解

国家自然科学基金

2+阅读 · 2013年12月31日

高阶Schwarz导数与Teichmuller空间紧化

国家自然科学基金

0+阅读 · 2012年12月31日

3DTV中编码效应消除和高质量双目视图生成研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于相机的低质量文本图像的复原与增强关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

局域高斯操作辅助下的连续变量量子纠缠蒸馏研究

国家自然科学基金

0+阅读 · 2012年12月31日

图像语义自动文本描述技术研究

国家自然科学基金

2+阅读 · 2012年12月31日

应用VIVI-OCT研究支架内血栓对支架内膜愈合模式影响及机制

国家自然科学基金

0+阅读 · 2011年12月31日

内窥镜手术中的三维实时图像导航关键技术

国家自然科学基金

0+阅读 · 2009年12月31日

图像处理问题的快速数值方法

国家自然科学基金

1+阅读 · 2008年12月31日

StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation

Arxiv

0+阅读 · 2023年5月31日

ProSpect: Expanded Conditioning for the Personalization of Attribute-aware Image Generation

Arxiv

0+阅读 · 2023年5月30日

A Federated Channel Modeling System using Generative Neural Networks

Arxiv

0+阅读 · 2023年5月30日

DiffSketching: Sketch Control Image Synthesis with Diffusion Models

Arxiv

0+阅读 · 2023年5月30日

RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths

Arxiv

0+阅读 · 2023年5月29日

GlyphControl: Glyph Conditional Control for Visual Text Generation

Arxiv

0+阅读 · 2023年5月29日

Transferring Visual Attributes from Natural Language to Verified Image Generation

Arxiv

0+阅读 · 2023年5月29日

FuseCap: Leveraging Large Language Models to Fuse Visual Data into Enriched Image Captions

Arxiv

0+阅读 · 2023年5月28日

GenerateCT: Text-Guided 3D Chest CT Generation

Arxiv

0+阅读 · 2023年5月26日

Deep Image Retrieval: A Survey

Arxiv

16+阅读 · 2021年1月27日

VIP会员

文章信息

相关主题

相关VIP内容

【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型

【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型

专知会员服务

22+阅读 · 2022年12月5日

中科院自动化所17篇CVPR 2022 新作速览！

中科院自动化所17篇CVPR 2022 新作速览！

专知会员服务

20+阅读 · 2022年3月19日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【百度&北京大学】自然语言生成的保真性:分析、评价和优化方法的系统综述，Faithfulness in Natural Language Generation: A Systematic Survey of Analysis, Evaluation and Optimization Methods

【百度&北京大学】自然语言生成的保真性:分析、评价和优化方法的系统综述，Faithfulness in Natural Language Generation: A Systematic Survey of Analysis, Evaluation and Optimization Methods

专知会员服务

15+阅读 · 2022年3月11日

【南洋理工大学Chuanxia Zheng博士论文】基于深度生成学习的逼真图像合成，197页pdf，Synthesizing Photorealistic Images with Deep Generative Learning

【南洋理工大学Chuanxia Zheng博士论文】基于深度生成学习的逼真图像合成，197页pdf，Synthesizing Photorealistic Images with Deep Generative Learning

专知会员服务

20+阅读 · 2022年3月9日

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

专知会员服务

23+阅读 · 2022年3月3日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知会员服务

77+阅读 · 2020年7月23日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

专知会员服务

33+阅读 · 2020年2月29日

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

专知会员服务

43+阅读 · 2020年1月28日

热门VIP内容

开通专知VIP会员享更多权益服务

新书册《几何深度学习的数学基础》

中程单向攻击无人机的战略意义：俄乌战争启示

在无标注条件下适配视觉—语言模型：全面综述

面向视觉语言模型的持续学习：遗忘之外的综述与分类体系

相关资讯

Stable Diffusion再迎重磅更新！2.0版「涩图」功能被砍，网友狂打差评

Stable Diffusion再迎重磅更新！2.0版「涩图」功能被砍，网友狂打差评

新智元

0+阅读 · 2022年11月25日

7 Papers & Radios | 谷歌推出DreamBooth扩散模型；张益唐零点猜想论文出炉

7 Papers & Radios | 谷歌推出DreamBooth扩散模型；张益唐零点猜想论文出炉

机器之心

2+阅读 · 2022年11月13日

谷歌新作Imagen：用Transformer和扩散模型把"文字到图像生成"卷上天！

谷歌新作Imagen：用Transformer和扩散模型把"文字到图像生成"卷上天！

CVer

0+阅读 · 2022年5月27日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知

15+阅读 · 2020年7月23日

【清华出品】NLP新方向文本对抗攻击与防御必读论文列表

【清华出品】NLP新方向文本对抗攻击与防御必读论文列表

专知

21+阅读 · 2019年7月11日

已删除

将门创投

12+阅读 · 2019年7月1日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

实战 | 用Python做图像处理（一）

实战 | 用Python做图像处理（一）

七月在线实验室

25+阅读 · 2018年5月23日

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

专知

15+阅读 · 2018年5月15日

Generative Adversarial Text to Image Synthesis论文解读

Generative Adversarial Text to Image Synthesis论文解读

统计学习与视觉计算组

13+阅读 · 2017年6月9日

相关论文

StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation

Arxiv

0+阅读 · 2023年5月31日

ProSpect: Expanded Conditioning for the Personalization of Attribute-aware Image Generation

Arxiv

0+阅读 · 2023年5月30日

A Federated Channel Modeling System using Generative Neural Networks

Arxiv

0+阅读 · 2023年5月30日

DiffSketching: Sketch Control Image Synthesis with Diffusion Models

Arxiv

0+阅读 · 2023年5月30日

RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths

Arxiv

0+阅读 · 2023年5月29日

GlyphControl: Glyph Conditional Control for Visual Text Generation

Arxiv

0+阅读 · 2023年5月29日

Transferring Visual Attributes from Natural Language to Verified Image Generation

Arxiv

0+阅读 · 2023年5月29日

FuseCap: Leveraging Large Language Models to Fuse Visual Data into Enriched Image Captions

Arxiv

0+阅读 · 2023年5月28日

GenerateCT: Text-Guided 3D Chest CT Generation

Arxiv

0+阅读 · 2023年5月26日

Deep Image Retrieval: A Survey

Arxiv

16+阅读 · 2021年1月27日

相关基金

血管内窥光声图像的频率谱研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于多任务概率视觉语义模型的图像场景理解

国家自然科学基金

2+阅读 · 2013年12月31日

高阶Schwarz导数与Teichmuller空间紧化

国家自然科学基金

0+阅读 · 2012年12月31日

3DTV中编码效应消除和高质量双目视图生成研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于相机的低质量文本图像的复原与增强关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

局域高斯操作辅助下的连续变量量子纠缠蒸馏研究

国家自然科学基金

0+阅读 · 2012年12月31日

图像语义自动文本描述技术研究

国家自然科学基金

2+阅读 · 2012年12月31日

应用VIVI-OCT研究支架内血栓对支架内膜愈合模式影响及机制

国家自然科学基金

0+阅读 · 2011年12月31日

内窥镜手术中的三维实时图像导航关键技术

国家自然科学基金

0+阅读 · 2009年12月31日

图像处理问题的快速数值方法

国家自然科学基金

1+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员