梦想艺术家:通过对抗性即时调试,走向可控的单制单制文本到图像的一代</s> (DreamArtist: Towards Controllable One-Shot Text-to-Image Generation via Contrastive Prompt-Tuning)

Large-scale text-to-image generation models have achieved remarkable progress in synthesizing high-quality, feature-rich images with high resolution guided by texts. However, these models often struggle with novel concepts, eg, new styles, object entities, etc. Although recent attempts have employed fine-tuning or prompt-tuning strategies to teach the pre-trained diffusion model novel concepts from a reference image set,they have the drawback of overfitting to the given reference images, particularly in one-shot applications, which is harmful to generate diverse and high-quality images while maintaining generation controllability. To tackle this challenge, we present a simple yet effective method called DreamArtist, which employs a positive-negative prompt-tuning learning strategy. Specifically, DreamArtist incorporates both positive and negative embeddings and jointly trains them. The positive embedding aggressively captures the salient characteristics of the reference image to drive diversified generation and the negative embedding rectifies inadequacies from the positive embedding. It learns not only what is correct, but also what can be avoided or improved. We have conducted extensive experiments and evaluated the proposed method from image similarity and diversity, generation controllability, and style cloning. And our DreamArtist has achieved a superior generation performance over existing methods. Besides, our additional evaluation on extended tasks, including concept compositions and prompt-guided image editing, demonstrates its effectiveness for more applications.

翻译：大规模文本到图像生成模型在以文本为指导,以高分辨率综合高品质、丰富地物图像方面取得了显著进展。然而,这些模型往往与新概念、例如新风格、新风格、物体实体等进行斗争。虽然最近尝试了微调或快速调整战略,从参考图像集中教授经过预先训练的传播模型新概念,但它们在过度适应特定参考图像方面有缺陷,特别是在一线应用中,这有害于生成多样化和高质量图像,同时保持生成控制能力。为了应对这一挑战,我们提出了一种简单而有效的方法,称为“梦想艺术”,采用了积极、消极的快速校正学习战略。具体地说,“梦想艺术”结合了积极和消极的调整战略,从参考图像集中深入地捕捉了参考图像的显著特征,以驱动多样化的生成,负面嵌入了从积极嵌入的缺陷。我们不仅学到了正确的东西,而且可以避免或改进什么。为了应对这一挑战,我们进行了广泛的实验,并评估了从图像相似性和更迅速调整学习的方法,并增加了我们制作的复制率和复制率,展示了我们制作方法。</s>

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

百篇论文纵览大型语言模型最新研究进展

专知会员服务

70+阅读 · 2023年3月31日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【因果基础】Causality Basics，36页ppt

专知会员服务

52+阅读 · 2021年8月8日

最新《Transformers模型》教程，64页ppt

专知会员服务

321+阅读 · 2020年11月26日