Form-NLU: 用于表格语言理解的数据集 (Form-NLU: Dataset for the Form Language Understanding) - 专知论文

会员服务 ·

0

信息提取 · NLU · 提取 · 数据集 · 结构 ·

2023 年 4 月 4 日

Form-NLU: Dataset for the Form Language Understanding

翻译：Form-NLU: 用于表格语言理解的数据集

Yihao Ding,Siqu Long,Jiabin Huang,Kaixuan Ren,Xingxiang Luo,Hyunsuk Chung,Soyeon Caren Han

from arxiv, Accepted by SIGIR 2023

Compared to general document analysis tasks, form document structure understanding and retrieval are challenging. Form documents are typically made by two types of authors; A form designer, who develops the form structure and keys, and a form user, who fills out form values based on the provided keys. Hence, the form values may not be aligned with the form designer's intention (structure and keys) if a form user gets confused. In this paper, we introduce Form-NLU, the first novel dataset for form structure understanding and its key and value information extraction, interpreting the form designer's intent and the alignment of user-written value on it. It consists of 857 form images, 6k form keys and values, and 4k table keys and values. Our dataset also includes three form types: digital, printed, and handwritten, which cover diverse form appearances and layouts. We propose a robust positional and logical relation-based form key-value information extraction framework. Using this dataset, Form-NLU, we first examine strong object detection models for the form layout understanding, then evaluate the key information extraction task on the dataset, providing fine-grained results for different types of forms and keys. Furthermore, we examine it with the off-the-shelf pdf layout extraction tool and prove its feasibility in real-world cases.

翻译：相比于一般的文档分析任务，表格文档结构的理解和检索更具挑战性。表格文档通常由两种类型的作者制作：一个是表单设计师，他开发表单结构和键值；另一个是表单用户，他根据提供的键值填写表单值。因此，如果表单用户困惑了，表单值可能与表单设计师的意图（结构和键值）不一致。在本文中，我们介绍了Form-NLU，这是一个用于表单结构理解及其键值信息提取的全新数据集，用于解释表单设计者的意图，并将用户编写的值与它对齐。它包括857张表格图像、6k 表格键和值和4k 表格键和值。我们的数据集还包括三种表格类型：数字、印刷和手写，涵盖了多种表格外观和布局。我们提出了一种强大的基于位置和逻辑关系的表格键值信息提取框架。使用这个数据集，我们首先检查了对于表格布局理解强大的物体检测模型，然后在数据集上评估了键信息提取任务，并为不同类型的表格和键提供了细粒度的结果。此外，我们使用现成的pdf布局提取工具对其进行了验证，并证明了其在实际案例中的可行性。

0

相关内容

信息提取

信息抽取也被称为事件抽取。与自动摘要相比,信息抽取更有目的性,并能将找到的信息以一定的框架展示。有时信息抽取也被用来完成自动摘要。

【2023新书】使用Python进行统计和数据可视化，554页pdf

【2023新书】使用Python进行统计和数据可视化，554页pdf

专知会员服务

130+阅读 · 2023年1月29日

【硬核书】基础架构作为代码、模式和实践:附带Python和terrform中的示例，402页pdf

【硬核书】基础架构作为代码、模式和实践:附带Python和terrform中的示例，402页pdf

专知会员服务

34+阅读 · 2022年8月24日

【2022新书】Transformer自然语言处理，Natural Language Processing with Transformers: Building Language Applications with Hugging Face

【2022新书】Transformer自然语言处理，Natural Language Processing with Transformers: Building Language Applications with Hugging Face

专知会员服务

522+阅读 · 2022年1月31日

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

专知会员服务

137+阅读 · 2021年11月6日

【ACL2020】Span-ConveRT：预训练对话表示小样本跨度提取，Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational Representations

【ACL2020】Span-ConveRT：预训练对话表示小样本跨度提取，Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational Representations

专知会员服务

17+阅读 · 2020年5月19日

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

专知会员服务

108+阅读 · 2020年5月1日

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

专知会员服务

92+阅读 · 2019年12月22日

【论文推荐】将机器语言模型扩展到人类级别的语言理解，Extending Machine Language Models toward Human-Level Language Understanding

【论文推荐】将机器语言模型扩展到人类级别的语言理解，Extending Machine Language Models toward Human-Level Language Understanding

专知会员服务

18+阅读 · 2019年12月14日

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

专知会员服务

20+阅读 · 2019年11月24日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

使用BERT做文本摘要

使用BERT做文本摘要

专知

23+阅读 · 2019年12月7日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Python自然语言处理: 使用SpaCycle库进行标记化、词干提取和词形还原

Python自然语言处理: 使用SpaCycle库进行标记化、词干提取和词形还原

Python程序员

18+阅读 · 2019年3月28日

时序数据异常检测工具/数据集大列表

时序数据异常检测工具/数据集大列表

极市平台

65+阅读 · 2019年2月23日

Github项目推荐 | Chatito - 使用简单的DSL为AI聊天机器人、NLP任务、命名实体识别或文本分类模型生成数据集

Github项目推荐 | Chatito - 使用简单的DSL为AI聊天机器人、NLP任务、命名实体识别或文本分类模型生成数据集

AI研习社

13+阅读 · 2019年1月21日

【PyTorch实战】手把手教你用torchtext处理文本数据

【PyTorch实战】手把手教你用torchtext处理文本数据

专知

13+阅读 · 2018年6月14日

干货 | 100+个NLP数据集大放送，再不愁数据！

干货 | 100+个NLP数据集大放送，再不愁数据！

数据派THU

11+阅读 · 2018年5月2日

自然语言处理（NLP）数据集整理

自然语言处理（NLP）数据集整理

论智

20+阅读 · 2018年4月8日

EB病毒miR-BART16调节凋亡信号通路促进淋巴瘤发生的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

面向实体信息集成的非合作半结构化深网数据源选择

国家自然科学基金

0+阅读 · 2014年12月31日

过渡金属双掺杂白光量子点的可控制备及LED应用

国家自然科学基金

0+阅读 · 2014年12月31日

miRNAs调控柿单宁合成代谢机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

新颖微纳结构界面调控的高效率柔性白光OLED器件

国家自然科学基金

0+阅读 · 2013年12月31日

基于Ontology的藏文语料库检索关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

面向ICN的可扩展命名数据路由机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于实例动态泛化的共指消解

国家自然科学基金

0+阅读 · 2009年12月31日

基于双路光相位调制光学倍频法的毫米波Radio Over Fiber系统研究

国家自然科学基金

0+阅读 · 2008年12月31日

面向复杂数据的生成器模式发现及其应用研究

国家自然科学基金

0+阅读 · 2008年12月31日

Personalized Dictionary Learning for Heterogeneous Datasets

Arxiv

0+阅读 · 2023年5月24日

On Degrees of Freedom in Defining and Testing Natural Language Understanding

Arxiv

0+阅读 · 2023年5月24日

Cream: Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

Arxiv

1+阅读 · 2023年5月24日

Large Language Models are Frame-level Directors for Zero-shot Text-to-Video Generation

Arxiv

0+阅读 · 2023年5月23日

Evaluation of African American Language Bias in Natural Language Generation

Arxiv

0+阅读 · 2023年5月23日

Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding

Arxiv

0+阅读 · 2023年5月22日

Understanding Diffusion Models: A Unified Perspective

Arxiv

14+阅读 · 2022年8月25日

LayoutLM: Pre-training of Text and Layout for Document Image Understanding

LayoutLM: Pre-training of Text and Layout for Document Image Understanding

Arxiv

12+阅读 · 2020年2月19日

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Arxiv

19+阅读 · 2020年2月15日

Transferring Common-Sense Knowledge for Object Detection

Arxiv

12+阅读 · 2018年4月3日

VIP会员

文章信息

相关主题

相关VIP内容

【2023新书】使用Python进行统计和数据可视化，554页pdf

【2023新书】使用Python进行统计和数据可视化，554页pdf

专知会员服务

130+阅读 · 2023年1月29日

【硬核书】基础架构作为代码、模式和实践:附带Python和terrform中的示例，402页pdf

【硬核书】基础架构作为代码、模式和实践:附带Python和terrform中的示例，402页pdf

专知会员服务

34+阅读 · 2022年8月24日

【2022新书】Transformer自然语言处理，Natural Language Processing with Transformers: Building Language Applications with Hugging Face

【2022新书】Transformer自然语言处理，Natural Language Processing with Transformers: Building Language Applications with Hugging Face

专知会员服务

522+阅读 · 2022年1月31日

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

专知会员服务

137+阅读 · 2021年11月6日

【ACL2020】Span-ConveRT：预训练对话表示小样本跨度提取，Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational Representations

【ACL2020】Span-ConveRT：预训练对话表示小样本跨度提取，Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational Representations

专知会员服务

17+阅读 · 2020年5月19日

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

专知会员服务

108+阅读 · 2020年5月1日

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

专知会员服务

92+阅读 · 2019年12月22日

【论文推荐】将机器语言模型扩展到人类级别的语言理解，Extending Machine Language Models toward Human-Level Language Understanding

【论文推荐】将机器语言模型扩展到人类级别的语言理解，Extending Machine Language Models toward Human-Level Language Understanding

专知会员服务

18+阅读 · 2019年12月14日

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

【NLP| 推荐文章】从统一文本到文本探讨迁移学习的局限性（Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer）

专知会员服务

20+阅读 · 2019年11月24日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

热门VIP内容

开通专知VIP会员享更多权益服务

【普林斯顿博士论文】在线学习：优化、控制与学习理论

不确定环境下无人机三维路径规划研究 | 221页

【NeurIPS2025】《LeapFactual：基于条件流匹配的可靠视觉反事实解释》

大语言模型将如何改变军事指挥结构

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

使用BERT做文本摘要

使用BERT做文本摘要

专知

23+阅读 · 2019年12月7日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Python自然语言处理: 使用SpaCycle库进行标记化、词干提取和词形还原

Python自然语言处理: 使用SpaCycle库进行标记化、词干提取和词形还原

Python程序员

18+阅读 · 2019年3月28日

时序数据异常检测工具/数据集大列表

时序数据异常检测工具/数据集大列表

极市平台

65+阅读 · 2019年2月23日

Github项目推荐 | Chatito - 使用简单的DSL为AI聊天机器人、NLP任务、命名实体识别或文本分类模型生成数据集

Github项目推荐 | Chatito - 使用简单的DSL为AI聊天机器人、NLP任务、命名实体识别或文本分类模型生成数据集

AI研习社

13+阅读 · 2019年1月21日

【PyTorch实战】手把手教你用torchtext处理文本数据

【PyTorch实战】手把手教你用torchtext处理文本数据

专知

13+阅读 · 2018年6月14日

干货 | 100+个NLP数据集大放送，再不愁数据！

干货 | 100+个NLP数据集大放送，再不愁数据！

数据派THU

11+阅读 · 2018年5月2日

自然语言处理（NLP）数据集整理

自然语言处理（NLP）数据集整理

论智

20+阅读 · 2018年4月8日

相关论文

Personalized Dictionary Learning for Heterogeneous Datasets

Arxiv

0+阅读 · 2023年5月24日

On Degrees of Freedom in Defining and Testing Natural Language Understanding

Arxiv

0+阅读 · 2023年5月24日

Cream: Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

Arxiv

1+阅读 · 2023年5月24日

Large Language Models are Frame-level Directors for Zero-shot Text-to-Video Generation

Arxiv

0+阅读 · 2023年5月23日

Evaluation of African American Language Bias in Natural Language Generation

Arxiv

0+阅读 · 2023年5月23日

Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding

Arxiv

0+阅读 · 2023年5月22日

Understanding Diffusion Models: A Unified Perspective

Arxiv

14+阅读 · 2022年8月25日

LayoutLM: Pre-training of Text and Layout for Document Image Understanding

LayoutLM: Pre-training of Text and Layout for Document Image Understanding

Arxiv

12+阅读 · 2020年2月19日

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Arxiv

19+阅读 · 2020年2月15日

Transferring Common-Sense Knowledge for Object Detection

Arxiv

12+阅读 · 2018年4月3日

相关基金

EB病毒miR-BART16调节凋亡信号通路促进淋巴瘤发生的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

面向实体信息集成的非合作半结构化深网数据源选择

国家自然科学基金

0+阅读 · 2014年12月31日

过渡金属双掺杂白光量子点的可控制备及LED应用

国家自然科学基金

0+阅读 · 2014年12月31日

miRNAs调控柿单宁合成代谢机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

新颖微纳结构界面调控的高效率柔性白光OLED器件

国家自然科学基金

0+阅读 · 2013年12月31日

基于Ontology的藏文语料库检索关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

面向ICN的可扩展命名数据路由机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于实例动态泛化的共指消解

国家自然科学基金

0+阅读 · 2009年12月31日

基于双路光相位调制光学倍频法的毫米波Radio Over Fiber系统研究

国家自然科学基金

0+阅读 · 2008年12月31日

面向复杂数据的生成器模式发现及其应用研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员