STEFANN：利用字体适应性神经网络的场景文本编辑器 (STEFANN: Scene Text Editor using Font Adaptive Neural Network)

Textual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our knowledge, there is no significant prior work that aims to modify the textual information in an image. The ability to edit text directly on images has several advantages including error correction, text restoration and image reusability. In this paper, we propose a method to modify text in an image at character-level. We approach the problem in two stages. At first, the unobserved character (target) is generated from an observed character (source) being modified. We propose two different neural network architectures - (a) FANnet to achieve structural consistency with source font and (b) Colornet to preserve source color. Next, we replace the source character with the generated character maintaining both geometric and visual consistency with neighboring characters. Our method works as a unified platform for modifying text in images. We present the effectiveness of our method on COCO-Text and ICDAR datasets both qualitatively and quantitatively.

翻译：摘要：捕捉到的场景中的文本信息对于场景的解释和决策具有重要作用。虽然存在能够成功检测和解释场景中复杂文本区域的方法，但据我们所知，以前不存在旨在修改图像中的文本信息的重要工作。直接在图像上编辑文本的能力具有几个优点，包括错误纠正、文本恢复和图像可重用性。在本文中，我们提出了一种在字符级别修改图像中的文本的方法。我们通过两个阶段来解决这个问题。首先，在观察到的正在修改的源字符中生成未观察到的字符（目标）。我们提出了两种不同的神经网络体系结构——（a）FANnet以实现与源字体的结构一致性和（b）Colornet以保持源颜色。接下来，我们用生成的字符替换源字符，并与相邻字符保持几何和视觉一致性。我们的方法作为修改图像中文本的统一平台。我们在COCO-Text和ICDAR数据集上以定性和定量的方式展示了我们方法的有效性。

相关内容

神经网络

关注 5910

人工神经网络（Artificial Neural Network，即ANN ），是20世纪80 年代以来人工智能领域兴起的研究热点。它从信息处理角度对人脑神经元网络进行抽象，建立某种简单模型，按不同的连接方式组成不同的网络。在工程与学术界也常直接简称为神经网络或类神经网络。神经网络是一种运算模型，由大量的节点（或称神经元）之间相互联接构成。每个节点代表一种特定的输出函数，称为激励函数（activation function）。每两个节点间的连接都代表一个对于通过该连接信号的加权值，称之为权重，这相当于人工神经网络的记忆。网络的输出则依网络的连接方式，权重值和激励函数的不同而不同。而网络自身通常都是对自然界某种算法或者函数的逼近，也可能是对一种逻辑策略的表达。最近十多年来，人工神经网络的研究工作不断深入，已经取得了很大的进展，其在模式识别、智能机器人、自动控制、预测估计、生物、医学、经济等领域已成功地解决了许多现代计算机难以解决的实际问题，表现出了良好的智能特性。

[ICCV 2021] 用于任意形状文本检测的自适应边界推荐网络

专知会员服务

11+阅读 · 2021年10月3日

【文献综述】Text Detection and Recognition in the Wild: A Review 自然文本检测与识别

专知会员服务

46+阅读 · 2020年6月11日

【CVPR2020】语义增强的场景文本识别的编码-解码器框架，SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

专知会员服务

25+阅读 · 2020年5月22日

【ACL2020】用于生成深度问题的语义图，Semantic Graphs for Generating Deep Questions

专知会员服务

26+阅读 · 2020年5月5日