多伙伴:用于识别复杂命名实体的大型多语言数据集 (MultiCoNER: A Large-scale Multilingual dataset for Complex Named Entity Recognition) - 专知论文

会员服务 ·

0

命名实体识别 · entity · 数据集 · MoDELS · 基准 ·

2022 年 8 月 30 日

MultiCoNER: A Large-scale Multilingual dataset for Complex Named Entity Recognition

翻译：多伙伴:用于识别复杂命名实体的大型多语言数据集

Shervin Malmasi,Anjie Fang,Besnik Fetahu,Sudipta Kar,Oleg Rokhlenko

from arxiv, Accepted at COLING 2022

We present MultiCoNER, a large multilingual dataset for Named Entity Recognition that covers 3 domains (Wiki sentences, questions, and search queries) across 11 languages, as well as multilingual and code-mixing subsets. This dataset is designed to represent contemporary challenges in NER, including low-context scenarios (short and uncased text), syntactically complex entities like movie titles, and long-tail entity distributions. The 26M token dataset is compiled from public resources using techniques such as heuristic-based sentence sampling, template extraction and slotting, and machine translation. We applied two NER models on our dataset: a baseline XLM-RoBERTa model, and a state-of-the-art GEMNET model that leverages gazetteers. The baseline achieves moderate performance (macro-F1=54%), highlighting the difficulty of our data. GEMNET, which uses gazetteers, improvement significantly (average improvement of macro-F1=+30%). MultiCoNER poses challenges even for large pre-trained language models, and we believe that it can help further research in building robust NER systems. MultiCoNER is publicly available at https://registry.opendata.aws/multiconer/ and we hope that this resource will help advance research in various aspects of NER.

翻译：我们为命名实体识别提供了多语种数据库,这是一个庞大的多语种数据库,涵盖11种语言的3个领域(维基句、问题和搜索查询)以及多语种和代码混合子集。该数据集旨在代表NER的当代挑战,包括低文本情景(短文本和未记录文本)、电影标题等综合复杂实体以及长尾实体分布。26M象征性数据集使用基于超自然的判刑抽样、模板抽取、插播和机器翻译等技术,从公共资源中汇编。我们在数据集上应用了两个NER模型:一个基线 XLM-ROBERTA模型,以及一个利用地名录的最先进的GEMNET模型。基准取得了中度的性能(macro-F1=54%),突出了我们数据的困难。GEMNET使用地名录,大大改进(宏观-F1 ⁇ 30% ) 。MultoneCON为大型预先培训语言模型提出了挑战,我们认为它能够帮助进一步研究强大的Opregreal NER系统。多COCONER将提供各种希望。

0

相关内容

命名实体识别

命名实体识别

命名实体识别（NER）（也称为实体标识，实体组块和实体提取）是信息抽取的子任务，旨在将非结构化文本中提到的命名实体定位和分类为预定义类别，例如人员姓名、地名、机构名、专有名词等。

知识荟萃

精品入门和进阶教程、论文和代码整理等

更多

查看相关VIP内容、论文、资讯等

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

50+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

【AAAI2020论文-清华大学】Enhanced Meta-Learning for Cross-lingual Named Entity Recognition with Minimal Resources，最小资源增强的元学习跨语言命名实体识别

【AAAI2020论文-清华大学】Enhanced Meta-Learning for Cross-lingual Named Entity Recognition with Minimal Resources，最小资源增强的元学习跨语言命名实体识别

专知会员服务

31+阅读 · 2019年11月17日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】图像分类必读开创性论文汇总

【推荐】图像分类必读开创性论文汇总

机器学习研究会

14+阅读 · 2017年8月15日

TLR4异常激活导致BM-MSCs衰老在SLE发生中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

新疆呼图壁地下储气库的GPS/InSAR监测及地质力学模拟研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于关键词的关系数据库查询技术研究

国家自然科学基金

0+阅读 · 2013年12月31日

半导体衬底上FeSe薄膜的外延生长及界面超导

国家自然科学基金

0+阅读 · 2013年12月31日

双曲平均曲率流

国家自然科学基金

0+阅读 · 2012年12月31日

杨桃根中DMDD基于TLR4靶标改善胰岛素抵抗的作用及分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

前列腺癌中Nedd4L对TrkA的抑癌性泛素化研究

国家自然科学基金

0+阅读 · 2012年12月31日

大型天文望远镜状态监控与故障诊断技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

灰绿霉素A的结构优化与抗肿瘤活性研究

国家自然科学基金

0+阅读 · 2009年12月31日

p75NTR对Alzheimer病Aβ20195;谢、沉积及其神经毒性作用的调控和机制

国家自然科学基金

0+阅读 · 2009年12月31日

STOP: A dataset for Spoken Task Oriented Semantic Parsing

STOP: A dataset for Spoken Task Oriented Semantic Parsing

Arxiv

0+阅读 · 2022年10月18日

RibSeg v2: A Large-scale Benchmark for Rib Labeling and Anatomical Centerline Extraction

Arxiv

0+阅读 · 2022年10月18日

SpanProto: A Two-stage Span-based Prototypical Network for Few-shot Named Entity Recognition

Arxiv

0+阅读 · 2022年10月17日

MV-HAN: A Hybrid Attentive Networks based Multi-View Learning Model for Large-scale Contents Recommendation

Arxiv

0+阅读 · 2022年10月14日

ConEntail: An Entailment-based Framework for Universal Zero and Few Shot Classification with Supervised Contrastive Pretraining

Arxiv

0+阅读 · 2022年10月14日

Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations

Arxiv

0+阅读 · 2022年10月14日

K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering

Arxiv

15+阅读 · 2021年9月22日

A Survey on Deep Learning for Named Entity Recognition

A Survey on Deep Learning for Named Entity Recognition

Arxiv

26+阅读 · 2020年3月13日

Knowledge-aware Graph Neural Networks with Label Smoothness Regularization for Recommendation

Arxiv

11+阅读 · 2019年6月13日

Incorporating Dictionaries into Deep Neural Networks for the Chinese Clinical Named Entity Recognition

Arxiv

12+阅读 · 2018年4月13日

VIP会员

文章信息

相关主题

命名实体识别

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

50+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

【AAAI2020论文-清华大学】Enhanced Meta-Learning for Cross-lingual Named Entity Recognition with Minimal Resources，最小资源增强的元学习跨语言命名实体识别

【AAAI2020论文-清华大学】Enhanced Meta-Learning for Cross-lingual Named Entity Recognition with Minimal Resources，最小资源增强的元学习跨语言命名实体识别

专知会员服务

31+阅读 · 2019年11月17日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军特种作战条令》最新102页

《洛克希德SR-71“黑鸟”侦察机动力系统》21页slides

美空军作战实验室通过人工智能和指挥控制技术创新推进杀伤链

《指挥控制能力分析方法论》最新报告

相关资讯

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】图像分类必读开创性论文汇总

【推荐】图像分类必读开创性论文汇总

机器学习研究会

14+阅读 · 2017年8月15日

相关论文

STOP: A dataset for Spoken Task Oriented Semantic Parsing

STOP: A dataset for Spoken Task Oriented Semantic Parsing

Arxiv

0+阅读 · 2022年10月18日

RibSeg v2: A Large-scale Benchmark for Rib Labeling and Anatomical Centerline Extraction

Arxiv

0+阅读 · 2022年10月18日

SpanProto: A Two-stage Span-based Prototypical Network for Few-shot Named Entity Recognition

Arxiv

0+阅读 · 2022年10月17日

MV-HAN: A Hybrid Attentive Networks based Multi-View Learning Model for Large-scale Contents Recommendation

Arxiv

0+阅读 · 2022年10月14日

ConEntail: An Entailment-based Framework for Universal Zero and Few Shot Classification with Supervised Contrastive Pretraining

Arxiv

0+阅读 · 2022年10月14日

Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations

Arxiv

0+阅读 · 2022年10月14日

K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering

Arxiv

15+阅读 · 2021年9月22日

A Survey on Deep Learning for Named Entity Recognition

A Survey on Deep Learning for Named Entity Recognition

Arxiv

26+阅读 · 2020年3月13日

Knowledge-aware Graph Neural Networks with Label Smoothness Regularization for Recommendation

Arxiv

11+阅读 · 2019年6月13日

Incorporating Dictionaries into Deep Neural Networks for the Chinese Clinical Named Entity Recognition

Arxiv

12+阅读 · 2018年4月13日

相关基金

TLR4异常激活导致BM-MSCs衰老在SLE发生中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

新疆呼图壁地下储气库的GPS/InSAR监测及地质力学模拟研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于关键词的关系数据库查询技术研究

国家自然科学基金

0+阅读 · 2013年12月31日

半导体衬底上FeSe薄膜的外延生长及界面超导

国家自然科学基金

0+阅读 · 2013年12月31日

双曲平均曲率流

国家自然科学基金

0+阅读 · 2012年12月31日

杨桃根中DMDD基于TLR4靶标改善胰岛素抵抗的作用及分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

前列腺癌中Nedd4L对TrkA的抑癌性泛素化研究

国家自然科学基金

0+阅读 · 2012年12月31日

大型天文望远镜状态监控与故障诊断技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

灰绿霉素A的结构优化与抗肿瘤活性研究

国家自然科学基金

0+阅读 · 2009年12月31日

p75NTR对Alzheimer病Aβ20195;谢、沉积及其神经毒性作用的调控和机制

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员