命名实体识别、多任务学习、内嵌实体、BERT、阿拉伯净净入学率公司 (Named Entity Recognition, Multi-Task Learning, Nested Entities, BERT, Arabic NER Corpus) - 专知论文

会员服务 ·

0

entity · 命名实体识别 · 多任务学习 · BERT · MoDELS ·

2022 年 5 月 19 日

Named Entity Recognition, Multi-Task Learning, Nested Entities, BERT, Arabic NER Corpus

翻译：命名实体识别、多任务学习、内嵌实体、BERT、阿拉伯净净入学率公司

Mustafa Jarrar,Mohammed Khalilia,Sana Ghanem

from arxiv, In Proceedings of the International Conference on Language Resources and Evaluation (LREC 2022), Marseille, France

This paper presents Wojood, a corpus for Arabic nested Named Entity Recognition (NER). Nested entities occur when one entity mention is embedded inside another entity mention. Wojood consists of about 550K Modern Standard Arabic (MSA) and dialect tokens that are manually annotated with 21 entity types including person, organization, location, event and date. More importantly, the corpus is annotated with nested entities instead of the more common flat annotations. The data contains about 75K entities and 22.5% of which are nested. The inter-annotator evaluation of the corpus demonstrated a strong agreement with Cohen's Kappa of 0.979 and an F1-score of 0.976. To validate our data, we used the corpus to train a nested NER model based on multi-task learning and AraBERT (Arabic BERT). The model achieved an overall micro F1-score of 0.884. Our corpus, the annotation guidelines, the source code and the pre-trained model are publicly available.

翻译：本文展示了Wojood, 阿拉伯嵌套命名实体识别(NER) 。当一个实体提到某个实体时, 就会出现堆积实体。 Wojood 由大约550K 现代阿拉伯文标准(MSA)和方言符号组成, 上面有21个实体类型, 包括个人、组织、地点、事件和日期, 手动加注, 更重要的是, 文中加注的是嵌套实体, 而不是更常见的平面说明。数据包含大约 75K 个实体, 其中22.5% 被嵌套。对堆积实体的评估显示, 与科恩的Kappa 的 0. 97979 和 F1 标码的 F1 0.976 达成强烈协议。为了验证我们的数据, 我们利用这个平台培训一个基于多任务学习和 AraBERT 的嵌嵌套净模型。该模型实现了0.884 总体的微型F1 核心。我们的体、指南、源代码和预培训模型可以公开查阅。

0

相关内容

entity

最新报告64页《军事中的人工智能和自主性：北约成员国的战略和部署概述》北约卓越合作网络防御中心，Artificial Intelligence and Autonomy in the Military: An Overview of NATO Member States’ Strategies and Deployment

最新报告64页《军事中的人工智能和自主性：北约成员国的战略和部署概述》北约卓越合作网络防御中心，Artificial Intelligence and Autonomy in the Military: An Overview of NATO Member States’ Strategies and Deployment

专知会员服务

30+阅读 · 2022年4月7日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【ACL2020】命名实体识别即依存解析，Named Entity Recognition as Dependency Parsing

【ACL2020】命名实体识别即依存解析，Named Entity Recognition as Dependency Parsing

专知会员服务

61+阅读 · 2020年5月15日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

专知

37+阅读 · 2018年2月21日

二维主体层板拓扑转变制备负载型金属催化剂介尺度表界面结构的构筑与调控

国家自然科学基金

0+阅读 · 2015年12月31日

面向ICN的网络级内嵌式缓存构架与配置管理方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

多功能金属纳米团簇的化学合成与组装

国家自然科学基金

2+阅读 · 2013年12月31日

VoIP流媒体隐写的在线检测模型与方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

采用原位同步辐射衍射研究纳米结构Cu/Ag多层膜的微机械行为

国家自然科学基金

0+阅读 · 2013年12月31日

面向中文指称概念的知识获取方法研究

国家自然科学基金

1+阅读 · 2012年12月31日

小尺寸超薄HfTiON/GGO堆栈高k栅介质InGaAs nMOSFET研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于多模态概率主题模型的实体相关文本可视化

国家自然科学基金

1+阅读 · 2011年12月31日

CH3-Si(111)基半导体外延结构的电沉积制备及光电性能研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于多层次语言粒度的文本情感分类研究

国家自然科学基金

1+阅读 · 2008年12月31日

AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition

AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition

Arxiv

0+阅读 · 2022年7月7日

Part-of-Speech Tagging of Odia Language Using statistical and Deep Learning-Based Approaches

Part-of-Speech Tagging of Odia Language Using statistical and Deep Learning-Based Approaches

Arxiv

0+阅读 · 2022年7月7日

Strong Heuristics for Named Entity Linking

Arxiv

0+阅读 · 2022年7月6日

Rethinking the Value of Gazetteer in Chinese Named Entity Recognition

Arxiv

1+阅读 · 2022年7月6日

CAN-NER: Convolutional Attention Network forChinese Named Entity Recognition

Arxiv

16+阅读 · 2019年4月3日

A Survey on Deep Learning for Named Entity Recognition

A Survey on Deep Learning for Named Entity Recognition

Arxiv

73+阅读 · 2018年12月22日

Incorporating Dictionaries into Deep Neural Networks for the Chinese Clinical Named Entity Recognition

Arxiv

12+阅读 · 2018年4月13日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

Deep Active Learning for Named Entity Recognition

Arxiv

15+阅读 · 2018年2月4日

Adversarial Learning for Chinese NER from Crowd Annotations

Arxiv

15+阅读 · 2018年1月16日

VIP会员

文章信息

相关主题

命名实体识别

多任务学习

相关VIP内容

最新报告64页《军事中的人工智能和自主性：北约成员国的战略和部署概述》北约卓越合作网络防御中心，Artificial Intelligence and Autonomy in the Military: An Overview of NATO Member States’ Strategies and Deployment

最新报告64页《军事中的人工智能和自主性：北约成员国的战略和部署概述》北约卓越合作网络防御中心，Artificial Intelligence and Autonomy in the Military: An Overview of NATO Member States’ Strategies and Deployment

专知会员服务

30+阅读 · 2022年4月7日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【ACL2020】命名实体识别即依存解析，Named Entity Recognition as Dependency Parsing

【ACL2020】命名实体识别即依存解析，Named Entity Recognition as Dependency Parsing

专知会员服务

61+阅读 · 2020年5月15日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

小规模训练指南：打造世界级大语言模型的关键方法

无人机编队飞行：复杂环境中作战的策略、挑战与应用

大模型APP，AI时代第一个爆款

从数据中心视角出发的高效大语言模型训练综述

相关资讯

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

专知

37+阅读 · 2018年2月21日

相关论文

AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition

AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition

Arxiv

0+阅读 · 2022年7月7日

Part-of-Speech Tagging of Odia Language Using statistical and Deep Learning-Based Approaches

Part-of-Speech Tagging of Odia Language Using statistical and Deep Learning-Based Approaches

Arxiv

0+阅读 · 2022年7月7日

Strong Heuristics for Named Entity Linking

Arxiv

0+阅读 · 2022年7月6日

Rethinking the Value of Gazetteer in Chinese Named Entity Recognition

Arxiv

1+阅读 · 2022年7月6日

CAN-NER: Convolutional Attention Network forChinese Named Entity Recognition

Arxiv

16+阅读 · 2019年4月3日

A Survey on Deep Learning for Named Entity Recognition

A Survey on Deep Learning for Named Entity Recognition

Arxiv

73+阅读 · 2018年12月22日

Incorporating Dictionaries into Deep Neural Networks for the Chinese Clinical Named Entity Recognition

Arxiv

12+阅读 · 2018年4月13日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

Deep Active Learning for Named Entity Recognition

Arxiv

15+阅读 · 2018年2月4日

Adversarial Learning for Chinese NER from Crowd Annotations

Arxiv

15+阅读 · 2018年1月16日

相关基金

二维主体层板拓扑转变制备负载型金属催化剂介尺度表界面结构的构筑与调控

国家自然科学基金

0+阅读 · 2015年12月31日

面向ICN的网络级内嵌式缓存构架与配置管理方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

多功能金属纳米团簇的化学合成与组装

国家自然科学基金

2+阅读 · 2013年12月31日

VoIP流媒体隐写的在线检测模型与方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

采用原位同步辐射衍射研究纳米结构Cu/Ag多层膜的微机械行为

国家自然科学基金

0+阅读 · 2013年12月31日

面向中文指称概念的知识获取方法研究

国家自然科学基金

1+阅读 · 2012年12月31日

小尺寸超薄HfTiON/GGO堆栈高k栅介质InGaAs nMOSFET研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于多模态概率主题模型的实体相关文本可视化

国家自然科学基金

1+阅读 · 2011年12月31日

CH3-Si(111)基半导体外延结构的电沉积制备及光电性能研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于多层次语言粒度的文本情感分类研究

国家自然科学基金

1+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员