MasaCtrl：一种无需调参的互相自注意力控制方法，实现一致的图像合成和编辑 (MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing) - 专知论文

会员服务 ·

0

非刚性 · 一致 · 调参 · 自注意力 · 图像生成 ·

2023 年 4 月 17 日

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

翻译：MasaCtrl：一种无需调参的互相自注意力控制方法，实现一致的图像合成和编辑

Mingdeng Cao,Xintao Wang,Zhongang Qi,Ying Shan,Xiaohu Qie,Yinqiang Zheng

from arxiv, Project available at https://ljzycmd.github.io/projects/MasaCtrl

Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex non-rigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing.

翻译：尽管大规模文本到图像生成和文本条件下的图像编辑已经取得了成功，但现有方法仍然难以产生一致的生成和编辑结果。例如，生成方法通常无法合成相同对象/字符的多个图像，但视角或姿态不同。与此同时，现有的编辑方法要么无法实现有效的复杂非刚性编辑，同时保持整体纹理和身份，要么需要耗费时间调整以捕捉特定于图像的外观。本文中，我们开发了MasaCtrl，这是一种无需调参的方法，可以同时实现一致的图像生成和复杂的非刚性图像编辑。具体来说，MasaCtrl将扩散模型中的现有自我注意力转换为相互自我注意力，以便查询源图像中的相关本地内容和纹理以实现一致性。为了进一步减轻前景和背景之间的查询混淆，我们提出了一种掩模引导的相互自注意力策略，其中掩模可以轻松地从交叉注意力图中提取。广泛的实验表明，所提出的MasaCtrl可以在一致的图像生成和复杂的非刚性真实图像编辑方面产生令人印象深刻的结果。

0

相关内容

非刚性

用GPT-4实现可控文本图像生成，UC伯克利&微软提出新框架Control-GPT

用GPT-4实现可控文本图像生成，UC伯克利&微软提出新框架Control-GPT

专知会员服务

35+阅读 · 2023年6月3日

【CVPR2023】基于图像特定提示学习的零样本生成模型自适应

【CVPR2023】基于图像特定提示学习的零样本生成模型自适应

专知会员服务

31+阅读 · 2023年4月7日

【Hugging Face】指导文本生成与约束波束搜索🤗Transformers，Guiding Text Generation with Constrained Beam Search in 🤗 Transformers

【Hugging Face】指导文本生成与约束波束搜索🤗Transformers，Guiding Text Generation with Constrained Beam Search in 🤗 Transformers

专知会员服务

22+阅读 · 2022年3月18日

【CVPR 2021】姿态可控的语音驱动说话人脸

专知会员服务

16+阅读 · 2021年5月13日

【ICML2020】统一预训练伪掩码语言模型

【ICML2020】统一预训练伪掩码语言模型

专知会员服务

27+阅读 · 2020年7月23日

【CVPR2020】通过自适应GANs生成不同的图像，Diverse Image Generation via Self-Conditioned GANs

【CVPR2020】通过自适应GANs生成不同的图像，Diverse Image Generation via Self-Conditioned GANs

专知会员服务

34+阅读 · 2020年6月19日

【CVPR2020】视觉跟踪的概率回归，Probabilistic Regression for Visual Tracking

【CVPR2020】视觉跟踪的概率回归，Probabilistic Regression for Visual Tracking

专知会员服务

37+阅读 · 2020年3月27日

【CVPR2020-Oral-牛津-Facebook】从单个图像进行端到端的视图合成，SynSin-View Synthesis

【CVPR2020-Oral-牛津-Facebook】从单个图像进行端到端的视图合成，SynSin-View Synthesis

专知会员服务

29+阅读 · 2020年3月26日

必读的10篇 CVPR 2019【生成对抗网络】相关论文和代码

必读的10篇 CVPR 2019【生成对抗网络】相关论文和代码

专知会员服务

33+阅读 · 2020年1月10日

【文章|自注意力(self-attention)机制图解】《Illustrated: Self-Attention》by Raimi Karim

【文章|自注意力(self-attention)机制图解】《Illustrated: Self-Attention》by Raimi Karim

专知会员服务

45+阅读 · 2019年11月18日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

【论文推荐】最新六篇视觉问答相关论文—深度嵌入学习、句子表征学习、深度特征聚合、3D匹配、细粒度文本摘要

【论文推荐】最新六篇视觉问答相关论文—深度嵌入学习、句子表征学习、深度特征聚合、3D匹配、细粒度文本摘要

专知

12+阅读 · 2018年6月9日

【论文推荐】最新四篇CVPR2018 视频描述生成相关论文—双向注意力、Transformer、重构网络、层次强化学习

【论文推荐】最新四篇CVPR2018 视频描述生成相关论文—双向注意力、Transformer、重构网络、层次强化学习

专知

31+阅读 · 2018年6月4日

Ian Goodfellow等提出自注意力GAN，ImageNet图像合成获最优结果！

Ian Goodfellow等提出自注意力GAN，ImageNet图像合成获最优结果！

新智元

11+阅读 · 2018年5月24日

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

专知

15+阅读 · 2018年5月15日

【论文推荐】最新八篇图像描述生成相关论文—比较级对抗学习、正则化RNNs、深层网络、视觉对话、婴儿说话、自我检索

【论文推荐】最新八篇图像描述生成相关论文—比较级对抗学习、正则化RNNs、深层网络、视觉对话、婴儿说话、自我检索

专知

10+阅读 · 2018年4月12日

【推荐】用Python/OpenCV实现增强现实

【推荐】用Python/OpenCV实现增强现实

机器学习研究会

15+阅读 · 2017年11月16日

液相费托合成反应的选择性调控新策略

国家自然科学基金

0+阅读 · 2012年12月31日

基于纯相位空间光调制器的近紫外飞秒多光束直写体布拉格光栅的研究

国家自然科学基金

0+阅读 · 2012年12月31日

半夏泻心汤调节2型糖尿病人GLP-1和β细胞功能的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

视觉表象可视化：基于个体脑激活模式重建视觉表象形象

国家自然科学基金

0+阅读 · 2012年12月31日

基于皮秒级精度可编程延时控制器的新型3D-TOF CMOS视觉传感系统关键问题的研究

国家自然科学基金

0+阅读 · 2012年12月31日

microRNA与转录因子不同基因型3'UTR的结合在先天性心脏病中的作用

国家自然科学基金

0+阅读 · 2011年12月31日

微分对策数值解法及非线性系统Min-Max鲁棒后退时域控制算法研究

国家自然科学基金

0+阅读 · 2009年12月31日

肾上腺源性及原发性高血压线粒体tRNAIle、tRNALeu(UUR)和tRNAlys基因突变的差异对比研究

国家自然科学基金

0+阅读 · 2009年12月31日

多张裁剪曲面拼接模型的水密融合

国家自然科学基金

0+阅读 · 2009年12月31日

TRPV1介导的神经肽释放在心肌梗死后炎症和凋亡中的调节作用

国家自然科学基金

0+阅读 · 2009年12月31日

"Let's not Quote out of Context": Unified Vision-Language Pretraining for Context Assisted Image Captioning

Arxiv

0+阅读 · 2023年6月1日

FDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion Models

Arxiv

0+阅读 · 2023年6月1日

In-Context Learning User Simulators for Task-Oriented Dialog Systems

Arxiv

0+阅读 · 2023年6月1日

Example-based Motion Synthesis via Generative Motion Matching

Arxiv

0+阅读 · 2023年6月1日

Understanding and Mitigating Copying in Diffusion Models

Arxiv

0+阅读 · 2023年5月31日

Exploring Regions of Interest: Visualizing Histological Image Classification for Breast Cancer using Deep Learning

Exploring Regions of Interest: Visualizing Histological Image Classification for Breast Cancer using Deep Learning

Arxiv

0+阅读 · 2023年5月31日

Learning Control by Iterative Inversion

Arxiv

0+阅读 · 2023年5月30日

DiffSketching: Sketch Control Image Synthesis with Diffusion Models

Arxiv

0+阅读 · 2023年5月30日

Deep Image Retrieval: A Survey

Arxiv

16+阅读 · 2021年1月27日

A survey on deep hashing for image retrieval

A survey on deep hashing for image retrieval

Arxiv

15+阅读 · 2020年6月10日

VIP会员

文章信息

相关主题

相关VIP内容

用GPT-4实现可控文本图像生成，UC伯克利&微软提出新框架Control-GPT

用GPT-4实现可控文本图像生成，UC伯克利&微软提出新框架Control-GPT

专知会员服务

35+阅读 · 2023年6月3日

【CVPR2023】基于图像特定提示学习的零样本生成模型自适应

【CVPR2023】基于图像特定提示学习的零样本生成模型自适应

专知会员服务

31+阅读 · 2023年4月7日

【Hugging Face】指导文本生成与约束波束搜索🤗Transformers，Guiding Text Generation with Constrained Beam Search in 🤗 Transformers

【Hugging Face】指导文本生成与约束波束搜索🤗Transformers，Guiding Text Generation with Constrained Beam Search in 🤗 Transformers

专知会员服务

22+阅读 · 2022年3月18日

【CVPR 2021】姿态可控的语音驱动说话人脸

专知会员服务

16+阅读 · 2021年5月13日

【ICML2020】统一预训练伪掩码语言模型

【ICML2020】统一预训练伪掩码语言模型

专知会员服务

27+阅读 · 2020年7月23日

【CVPR2020】通过自适应GANs生成不同的图像，Diverse Image Generation via Self-Conditioned GANs

【CVPR2020】通过自适应GANs生成不同的图像，Diverse Image Generation via Self-Conditioned GANs

专知会员服务

34+阅读 · 2020年6月19日

【CVPR2020】视觉跟踪的概率回归，Probabilistic Regression for Visual Tracking

【CVPR2020】视觉跟踪的概率回归，Probabilistic Regression for Visual Tracking

专知会员服务

37+阅读 · 2020年3月27日

【CVPR2020-Oral-牛津-Facebook】从单个图像进行端到端的视图合成，SynSin-View Synthesis

【CVPR2020-Oral-牛津-Facebook】从单个图像进行端到端的视图合成，SynSin-View Synthesis

专知会员服务

29+阅读 · 2020年3月26日

必读的10篇 CVPR 2019【生成对抗网络】相关论文和代码

必读的10篇 CVPR 2019【生成对抗网络】相关论文和代码

专知会员服务

33+阅读 · 2020年1月10日

【文章|自注意力(self-attention)机制图解】《Illustrated: Self-Attention》by Raimi Karim

【文章|自注意力(self-attention)机制图解】《Illustrated: Self-Attention》by Raimi Karim

专知会员服务

45+阅读 · 2019年11月18日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军特种作战条令》最新102页

《洛克希德SR-71“黑鸟”侦察机动力系统》21页slides

美空军作战实验室通过人工智能和指挥控制技术创新推进杀伤链

《指挥控制能力分析方法论》最新报告

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

【论文推荐】最新六篇视觉问答相关论文—深度嵌入学习、句子表征学习、深度特征聚合、3D匹配、细粒度文本摘要

【论文推荐】最新六篇视觉问答相关论文—深度嵌入学习、句子表征学习、深度特征聚合、3D匹配、细粒度文本摘要

专知

12+阅读 · 2018年6月9日

【论文推荐】最新四篇CVPR2018 视频描述生成相关论文—双向注意力、Transformer、重构网络、层次强化学习

【论文推荐】最新四篇CVPR2018 视频描述生成相关论文—双向注意力、Transformer、重构网络、层次强化学习

专知

31+阅读 · 2018年6月4日

Ian Goodfellow等提出自注意力GAN，ImageNet图像合成获最优结果！

Ian Goodfellow等提出自注意力GAN，ImageNet图像合成获最优结果！

新智元

11+阅读 · 2018年5月24日

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

专知

15+阅读 · 2018年5月15日

【论文推荐】最新八篇图像描述生成相关论文—比较级对抗学习、正则化RNNs、深层网络、视觉对话、婴儿说话、自我检索

【论文推荐】最新八篇图像描述生成相关论文—比较级对抗学习、正则化RNNs、深层网络、视觉对话、婴儿说话、自我检索

专知

10+阅读 · 2018年4月12日

【推荐】用Python/OpenCV实现增强现实

【推荐】用Python/OpenCV实现增强现实

机器学习研究会

15+阅读 · 2017年11月16日

相关论文

"Let's not Quote out of Context": Unified Vision-Language Pretraining for Context Assisted Image Captioning

Arxiv

0+阅读 · 2023年6月1日

FDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion Models

Arxiv

0+阅读 · 2023年6月1日

In-Context Learning User Simulators for Task-Oriented Dialog Systems

Arxiv

0+阅读 · 2023年6月1日

Example-based Motion Synthesis via Generative Motion Matching

Arxiv

0+阅读 · 2023年6月1日

Understanding and Mitigating Copying in Diffusion Models

Arxiv

0+阅读 · 2023年5月31日

Exploring Regions of Interest: Visualizing Histological Image Classification for Breast Cancer using Deep Learning

Exploring Regions of Interest: Visualizing Histological Image Classification for Breast Cancer using Deep Learning

Arxiv

0+阅读 · 2023年5月31日

Learning Control by Iterative Inversion

Arxiv

0+阅读 · 2023年5月30日

DiffSketching: Sketch Control Image Synthesis with Diffusion Models

Arxiv

0+阅读 · 2023年5月30日

Deep Image Retrieval: A Survey

Arxiv

16+阅读 · 2021年1月27日

A survey on deep hashing for image retrieval

A survey on deep hashing for image retrieval

Arxiv

15+阅读 · 2020年6月10日

相关基金

液相费托合成反应的选择性调控新策略

国家自然科学基金

0+阅读 · 2012年12月31日

基于纯相位空间光调制器的近紫外飞秒多光束直写体布拉格光栅的研究

国家自然科学基金

0+阅读 · 2012年12月31日

半夏泻心汤调节2型糖尿病人GLP-1和β细胞功能的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

视觉表象可视化：基于个体脑激活模式重建视觉表象形象

国家自然科学基金

0+阅读 · 2012年12月31日

基于皮秒级精度可编程延时控制器的新型3D-TOF CMOS视觉传感系统关键问题的研究

国家自然科学基金

0+阅读 · 2012年12月31日

microRNA与转录因子不同基因型3'UTR的结合在先天性心脏病中的作用

国家自然科学基金

0+阅读 · 2011年12月31日

微分对策数值解法及非线性系统Min-Max鲁棒后退时域控制算法研究

国家自然科学基金

0+阅读 · 2009年12月31日

肾上腺源性及原发性高血压线粒体tRNAIle、tRNALeu(UUR)和tRNAlys基因突变的差异对比研究

国家自然科学基金

0+阅读 · 2009年12月31日

多张裁剪曲面拼接模型的水密融合

国家自然科学基金

0+阅读 · 2009年12月31日

TRPV1介导的神经肽释放在心肌梗死后炎症和凋亡中的调节作用

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员