场景风格文本编辑 (Scene Style Text Editing) - 专知论文

会员服务 ·

0

特征空间 · 潜在 · 生成器 · 嵌入 · 融合 ·

2023 年 4 月 20 日

Scene Style Text Editing

翻译：场景风格文本编辑

Tonghua Su,Fuxiang Yang,Xiang Zhou,Donglin Di,Zhongjie Wang,Songze Li

In this work, we propose a task called "Scene Style Text Editing (SSTE)", changing the text content as well as the text style of the source image while keeping the original text scene. Existing methods neglect to fine-grained adjust the style of the foreground text, such as its rotation angle, color, and font type. To tackle this task, we propose a quadruple framework named "QuadNet" to embed and adjust foreground text styles in the latent feature space. Specifically, QuadNet consists of four parts, namely background inpainting, style encoder, content encoder, and fusion generator. The background inpainting erases the source text content and recovers the appropriate background with a highly authentic texture. The style encoder extracts the style embedding of the foreground text. The content encoder provides target text representations in the latent feature space to implement the content edits. The fusion generator combines the information yielded from the mentioned parts and generates the rendered text images. Practically, our method is capable of performing promisingly on real-world datasets with merely string-level annotation. To the best of our knowledge, our work is the first to finely manipulate the foreground text content and style by deeply semantic editing in the latent feature space. Extensive experiments demonstrate that QuadNet has the ability to generate photo-realistic foreground text and avoid source text shadows in real-world scenes when editing text content.

翻译：在本文中，我们提出了一项名为“场景风格文本编辑（SSTE）”的任务，该任务旨在在保持原始文本场景的情况下更改源图像的文本内容和文本样式。现有方法忽略调整前景文本的风格，例如旋转角度、颜色和字体类型等方面。为了解决这个问题，我们提出了一个名称为“ QuadNet”的四重框架，以在潜在特征空间中嵌入和调整前景文本样式。具体而言，QuadNet由四个部分组成，即背景修复、样式编码器、内容编码器和融合生成器。背景修复擦除源文本内容并恢复具有高度真实质感的背景。样式编码器提取前景文本的样式嵌入。内容编码器在潜在特征空间中提供目标文本表示，以实现内容编辑。融合生成器将从上述部分得到的信息结合起来，并生成渲染文本图像。在实际应用中，我们的方法能够在仅具有字符串级注释的情况下，在真实世界的数据集上表现出色。据我们所知，我们的工作是第一个通过对潜在特征空间进行深度语义编辑，精细地操纵前景文本内容和样式的工作。广泛的实验表明，QuadNet能够生成逼真的前景文本，并在编辑文本内容时避免源文本阴影。

0

相关内容

特征空间

【2022新书】文本生成的深度学习方法，201页pdf，Deep Learning Approaches to Text Production

【2022新书】文本生成的深度学习方法，201页pdf，Deep Learning Approaches to Text Production

专知会员服务

39+阅读 · 2022年5月28日

【CVPR 2022】基于Transformer的图象风格化，StyTr2: Image Style Transfer with Transformers

【CVPR 2022】基于Transformer的图象风格化，StyTr2: Image Style Transfer with Transformers

专知会员服务

11+阅读 · 2022年3月19日

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

专知会员服务

27+阅读 · 2022年3月3日

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

专知会员服务

23+阅读 · 2022年3月3日

【ICCV 2021】HCFlow：使用一个统一的框架处理图像超分辨率和图像再缩放

专知会员服务

15+阅读 · 2021年10月4日

【CVPR2020】语义增强的场景文本识别的编码-解码器框架，SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

【CVPR2020】语义增强的场景文本识别的编码-解码器框架，SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

专知会员服务

25+阅读 · 2020年5月22日

【香港中文大学-CVPR2020】Rotate-and-Render: Unsupervised Photorealistic Face Rotation from Single-View Images

【香港中文大学-CVPR2020】Rotate-and-Render: Unsupervised Photorealistic Face Rotation from Single-View Images

专知会员服务

22+阅读 · 2020年3月18日

微软亚洲研究院新论文-《多模态预训练语言模型UniViLM》面向多模态理解和生成的统一视频和语言预训练模型

微软亚洲研究院新论文-《多模态预训练语言模型UniViLM》面向多模态理解和生成的统一视频和语言预训练模型

专知会员服务

109+阅读 · 2020年2月19日

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

专知会员服务

20+阅读 · 2019年11月24日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

SIGGRAPH Asia 2022 | 一句话生成高清360度场景及光照，可直接渲染数字资产

SIGGRAPH Asia 2022 | 一句话生成高清360度场景及光照，可直接渲染数字资产

机器之心

0+阅读 · 2022年10月5日

7 Papers & Radios | 国产数据库入选顶会VLDB 2022；一句话生成高清360度场景和光照

7 Papers & Radios | 国产数据库入选顶会VLDB 2022；一句话生成高清360度场景和光照

机器之心

0+阅读 · 2022年10月2日

|[IEEE TIP 2020]EraseNet：端到端的真实场景文本擦除方法

|[IEEE TIP 2020]EraseNet：端到端的真实场景文本擦除方法

专知

18+阅读 · 2020年10月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

图像和文本的融合表示学习——Text2Image和Image2Text

图像和文本的融合表示学习——Text2Image和Image2Text

专知

125+阅读 · 2018年6月11日

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

专知

15+阅读 · 2018年5月15日

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

专知

15+阅读 · 2018年2月13日

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

专知

66+阅读 · 2018年1月31日

Generative Adversarial Text to Image Synthesis论文解读

Generative Adversarial Text to Image Synthesis论文解读

统计学习与视觉计算组

13+阅读 · 2017年6月9日

基于多源语义表示学习的社交媒体文本属性情感分类研究

国家自然科学基金

4+阅读 · 2017年12月31日

基于短文本的知识库自动更新关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于语谱图信息的汉语词汇整体识别和语音增强方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

自由视点人体活动识别中的稀疏表达与学习

国家自然科学基金

0+阅读 · 2013年12月31日

基于图像的室外场景光影分析与编辑

国家自然科学基金

0+阅读 · 2013年12月31日

基于Ontology的藏文语料库检索关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于局部特征的自然场景下文字定位和识别研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于增强现实的精确截骨手术导航系统

国家自然科学基金

1+阅读 · 2012年12月31日

嵌入式系统构件模型的领域语义检查方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

高精细模型的向量位移映射表示及几何处理

国家自然科学基金

0+阅读 · 2011年12月31日

Stable Diffusion is Unstable

Arxiv

0+阅读 · 2023年6月6日

QuantArt: Quantizing Image Style Transfer Towards High Visual Fidelity

Arxiv

0+阅读 · 2023年6月5日

HeadSculpt: Crafting 3D Head Avatars with Text

Arxiv

0+阅读 · 2023年6月5日

Stable Diffusion is Untable

Arxiv

0+阅读 · 2023年6月5日

AvatarStudio: Text-driven Editing of 3D Dynamic Human Head Avatars

Arxiv

0+阅读 · 2023年6月2日

Text Style Transfer Back-Translation

Arxiv

0+阅读 · 2023年6月2日

CLIP-Layout: Style-Consistent Indoor Scene Synthesis with Semantic Furniture Embedding

Arxiv

0+阅读 · 2023年6月2日

NeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Models

Arxiv

42+阅读 · 2023年4月19日

Question-controlled Text-aware Image Captioning

Arxiv

10+阅读 · 2021年8月4日

Rotation-Sensitive Regression for Oriented Scene Text Detection

Arxiv

13+阅读 · 2018年3月14日

VIP会员

文章信息

相关主题

相关VIP内容

【2022新书】文本生成的深度学习方法，201页pdf，Deep Learning Approaches to Text Production

【2022新书】文本生成的深度学习方法，201页pdf，Deep Learning Approaches to Text Production

专知会员服务

39+阅读 · 2022年5月28日

【CVPR 2022】基于Transformer的图象风格化，StyTr2: Image Style Transfer with Transformers

【CVPR 2022】基于Transformer的图象风格化，StyTr2: Image Style Transfer with Transformers

专知会员服务

11+阅读 · 2022年3月19日

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

专知会员服务

27+阅读 · 2022年3月3日

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

【CVPR 2022】可控图像合成与编辑的合成生成先验学习，SemanticStyleGAN: Learning Compositonal Generative Priors for Controllable Image Synthesis and Editing

专知会员服务

23+阅读 · 2022年3月3日

【ICCV 2021】HCFlow：使用一个统一的框架处理图像超分辨率和图像再缩放

专知会员服务

15+阅读 · 2021年10月4日

【CVPR2020】语义增强的场景文本识别的编码-解码器框架，SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

【CVPR2020】语义增强的场景文本识别的编码-解码器框架，SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

专知会员服务

25+阅读 · 2020年5月22日

【香港中文大学-CVPR2020】Rotate-and-Render: Unsupervised Photorealistic Face Rotation from Single-View Images

【香港中文大学-CVPR2020】Rotate-and-Render: Unsupervised Photorealistic Face Rotation from Single-View Images

专知会员服务

22+阅读 · 2020年3月18日

微软亚洲研究院新论文-《多模态预训练语言模型UniViLM》面向多模态理解和生成的统一视频和语言预训练模型

微软亚洲研究院新论文-《多模态预训练语言模型UniViLM》面向多模态理解和生成的统一视频和语言预训练模型

专知会员服务

109+阅读 · 2020年2月19日

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

专知会员服务

20+阅读 · 2019年11月24日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《城市滨海地区：理解复杂多变环境下的指挥控制框架》50页报告

《理解城市战及其在俄乌战争中的表现》报告

美空军“顶点2025”实验：推进AI在C2、动态目标锁定与联盟集成中的应用

《建设式兵棋模拟作为战术集群配置优化的关键组成部分》

相关资讯

SIGGRAPH Asia 2022 | 一句话生成高清360度场景及光照，可直接渲染数字资产

SIGGRAPH Asia 2022 | 一句话生成高清360度场景及光照，可直接渲染数字资产

机器之心

0+阅读 · 2022年10月5日

7 Papers & Radios | 国产数据库入选顶会VLDB 2022；一句话生成高清360度场景和光照

7 Papers & Radios | 国产数据库入选顶会VLDB 2022；一句话生成高清360度场景和光照

机器之心

0+阅读 · 2022年10月2日

|[IEEE TIP 2020]EraseNet：端到端的真实场景文本擦除方法

|[IEEE TIP 2020]EraseNet：端到端的真实场景文本擦除方法

专知

18+阅读 · 2020年10月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

图像和文本的融合表示学习——Text2Image和Image2Text

图像和文本的融合表示学习——Text2Image和Image2Text

专知

125+阅读 · 2018年6月11日

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

【论文推荐】最新八篇生成对抗网络相关论文—条件翻译、RGB-D动作识别、量子生成对抗网络、语义对齐、视频摘要、视觉-文本注意力

专知

15+阅读 · 2018年5月15日

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

专知

15+阅读 · 2018年2月13日

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

专知

66+阅读 · 2018年1月31日

Generative Adversarial Text to Image Synthesis论文解读

Generative Adversarial Text to Image Synthesis论文解读

统计学习与视觉计算组

13+阅读 · 2017年6月9日

相关论文

Stable Diffusion is Unstable

Arxiv

0+阅读 · 2023年6月6日

QuantArt: Quantizing Image Style Transfer Towards High Visual Fidelity

Arxiv

0+阅读 · 2023年6月5日

HeadSculpt: Crafting 3D Head Avatars with Text

Arxiv

0+阅读 · 2023年6月5日

Stable Diffusion is Untable

Arxiv

0+阅读 · 2023年6月5日

AvatarStudio: Text-driven Editing of 3D Dynamic Human Head Avatars

Arxiv

0+阅读 · 2023年6月2日

Text Style Transfer Back-Translation

Arxiv

0+阅读 · 2023年6月2日

CLIP-Layout: Style-Consistent Indoor Scene Synthesis with Semantic Furniture Embedding

Arxiv

0+阅读 · 2023年6月2日

NeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Models

Arxiv

42+阅读 · 2023年4月19日

Question-controlled Text-aware Image Captioning

Arxiv

10+阅读 · 2021年8月4日

Rotation-Sensitive Regression for Oriented Scene Text Detection

Arxiv

13+阅读 · 2018年3月14日

相关基金

基于多源语义表示学习的社交媒体文本属性情感分类研究

国家自然科学基金

4+阅读 · 2017年12月31日

基于短文本的知识库自动更新关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于语谱图信息的汉语词汇整体识别和语音增强方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

自由视点人体活动识别中的稀疏表达与学习

国家自然科学基金

0+阅读 · 2013年12月31日

基于图像的室外场景光影分析与编辑

国家自然科学基金

0+阅读 · 2013年12月31日

基于Ontology的藏文语料库检索关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于局部特征的自然场景下文字定位和识别研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于增强现实的精确截骨手术导航系统

国家自然科学基金

1+阅读 · 2012年12月31日

嵌入式系统构件模型的领域语义检查方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

高精细模型的向量位移映射表示及几何处理

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员