This paper presents results of Document Visual Question Answering Challenge organized as part of "Text and Documents in the Deep Learning Era" workshop, in CVPR 2020. The challenge introduces a new problem - Visual Question Answering on document images. The challenge comprised two tasks. The first task concerns with asking questions on a single document image. On the other hand, the second task is set as a retrieval task where the question is posed over a collection of images. For the task 1 a new dataset is introduced comprising 50,000 questions-answer(s) pairs defined over 12,767 document images. For task 2 another dataset has been created comprising 20 questions over 14,362 document images which share the same document template.
翻译:本文件介绍了作为2020年CVPR CVPR“深学习时代的文本和文件”讲习班的一部分而组织的文档视觉回答挑战的结果。 挑战提出了一个新问题—— 文件图像的视觉回答。 挑战包括两个任务。 第一个任务涉及在单个文档图像上询问问题。 另一方面, 第二个任务被设定为对图像收集提出问题的检索任务。 对于任务1, 引入了一个新的数据集, 由超过12,767个文件图像定义的50,000对问答对组成。 对于任务2, 创建了另一个数据集, 由超过14,362个文件图像的20个问题组成, 共享相同的文档模板 。