低资源的日尔曼语系方言和语言的语料库综述 (A Survey of Corpora for Germanic Low-Resource Languages and Dialects) - 专知论文

会员服务 ·

0

低资源 · 语料库 · 语料 · NLP · 注释（编程） ·

2023 年 4 月 19 日

A Survey of Corpora for Germanic Low-Resource Languages and Dialects

翻译：低资源的日尔曼语系方言和语言的语料库综述

Verena Blaschke,Hinrich Schütze,Barbara Plank

from arxiv, NoDaLiDa 2023

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized low-resource languages. Even within branches of major language families, often considered well-researched, little is known about the extent and type of available resources and what the major NLP challenges are for these language varieties. The first step to address this situation is a systematic survey of available corpora (most importantly, annotated corpora, which are particularly valuable for NLP research). Focusing on Germanic low-resource language varieties, we provide such a survey in this paper. Except for geolocation (origin of speaker or document), we find that manually annotated linguistic resources are sparse and, if they exist, mostly cover morphosyntax. Despite this lack of resources, we observe that interest in this area is increasing: there is active development and a growing research community. To facilitate research, we make our overview of over 80 corpora publicly available. We share a companion website of this overview at https://github.com/mainlp/germanic-lrl-corpora .

翻译：尽管近年来取得了很大的进展，自然语言处理（NLP）的绝大部分工作仍是针对具有许多说话者的标准语言。在本文中，我们转而关注低资源语言，特别是非标准化的低资源语言。即使在被认为已经有很多研究的主要语系中，这些语言的资源可用性、类型及其NLP研究的主要挑战仍鲜为人知。解决这种情况的第一步是对可用语料库进行系统的调查（最重要的是手动注释的语料库，对于NLP研究尤其有价值）。在本文中，我们将重点放在日尔曼语族的低资源语言上，提供这样的调查。除了地理位置（说话者或文档的起源）之外，我们发现手动注释的语言资源很少，如果存在的话，主要是涵盖词法和句法。尽管这种资源匮乏，我们观察到对该领域的兴趣正在增加：有着积极的发展和不断增长的研究社区。为了促进研究，我们公开了80多个语料库的概述。我们在 https://github.com/mainlp/germanic-lrl-corpora 上分享了这个概述的陪伴网站。

0

相关内容

低资源

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

专知会员服务

140+阅读 · 2020年7月10日

【哈工大】基于文档的对话系统(DGDS)综述，A Survey of Document Grounded Dialogue Systems (DGDS)

【哈工大】基于文档的对话系统(DGDS)综述，A Survey of Document Grounded Dialogue Systems (DGDS)

专知会员服务

35+阅读 · 2020年4月30日

【CMU-TACL2020】低资源跨语言实体链接，Low-resource Crosslingual EntityLinking

专知会员服务

17+阅读 · 2020年3月29日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【论文】多语言神经机器翻译综述（A Comprehensive Survey of Multilingual Neural Machine Translation）

【论文】多语言神经机器翻译综述（A Comprehensive Survey of Multilingual Neural Machine Translation）

专知会员服务

20+阅读 · 2020年1月7日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

上百种预训练中文词向量：Chinese-Word-Vectors

上百种预训练中文词向量：Chinese-Word-Vectors

AINLP

23+阅读 · 2019年2月26日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

专知

37+阅读 · 2018年2月21日

【论文推荐】最新5篇信息抽取（IE）相关论文—开放信息抽取、不完整信息、主动学习、越南语、依存分析

【论文推荐】最新5篇信息抽取（IE）相关论文—开放信息抽取、不完整信息、主动学习、越南语、依存分析

专知

12+阅读 · 2018年2月2日

【论文推荐】最新5篇情感分析相关论文—深度学习情感分析综述、情感分析语料库、情感预测性、上下文和位置感知的因子分解模型、LSTM

【论文推荐】最新5篇情感分析相关论文—深度学习情感分析综述、情感分析语料库、情感预测性、上下文和位置感知的因子分解模型、LSTM

专知

55+阅读 · 2018年1月28日

【推荐】深度学习情感分析综述

【推荐】深度学习情感分析综述

机器学习研究会

58+阅读 · 2018年1月26日

新疆地区维吾尔族、汉族klotho基因单核苷酸多态性与特发性草酸钙肾结石相关性研究

国家自然科学基金

0+阅读 · 2014年12月31日

钬、镨双掺杂氟化镥锂晶体2.9μm中红外激光特性研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于紫外光谱的太阳耀斑能量和动力学过程研究

国家自然科学基金

0+阅读 · 2012年12月31日

CXCR7/SDF-1/ITAC信号调控前列腺癌细胞迁徙、侵袭及增殖的作用机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

复域差分, 差分方程和微分方程的研究

国家自然科学基金

0+阅读 · 2011年12月31日

中文自动口语摘要技术研究

国家自然科学基金

1+阅读 · 2011年12月31日

家庭高等教育投资行为实证研究

国家自然科学基金

0+阅读 · 2011年12月31日

可调谐拍长线扫描的不连续表面三维干涉测量研究

国家自然科学基金

0+阅读 · 2009年12月31日

赖氨酸特异性去甲基酶1对前列腺癌雄激素非依赖性进展的影响及机制

国家自然科学基金

0+阅读 · 2009年12月31日

三波段双极化共用口径SAR天线阵的研究

国家自然科学基金

0+阅读 · 2008年12月31日

Weakly-Supervised Conditional Embedding for Referred Visual Search

Arxiv

0+阅读 · 2023年6月5日

The State of the Art in Creating Visualization Corpora for Automated Chart Analysis

Arxiv

0+阅读 · 2023年6月4日

UCAS-IIE-NLP at SemEval-2023 Task 12: Enhancing Generalization of Multilingual BERT for Low-resource Sentiment Analysis

Arxiv

0+阅读 · 2023年6月1日

Neural Natural Language Processing for Long Texts: A Survey of the State-of-the-Art

Arxiv

0+阅读 · 2023年6月1日

BeamSearchQA: Large Language Models are Strong Zero-Shot QA Solver

Arxiv

0+阅读 · 2023年6月1日

Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili

Arxiv

0+阅读 · 2023年6月1日

Meta Learning for Natural Language Processing: A Survey

Meta Learning for Natural Language Processing: A Survey

Arxiv

14+阅读 · 2022年5月3日

Pretrained Transformers for Text Ranking: BERT and Beyond

Arxiv

28+阅读 · 2020年10月13日

Multilingual Sentiment Analysis: An RNN-Based Framework for Limited Data

Arxiv

12+阅读 · 2018年6月8日

EventKG: A Multilingual Event-Centric Temporal Knowledge Graph

Arxiv

11+阅读 · 2018年4月12日

VIP会员

文章信息

相关主题

注释（编程）

相关VIP内容

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

专知会员服务

140+阅读 · 2020年7月10日

【哈工大】基于文档的对话系统(DGDS)综述，A Survey of Document Grounded Dialogue Systems (DGDS)

【哈工大】基于文档的对话系统(DGDS)综述，A Survey of Document Grounded Dialogue Systems (DGDS)

专知会员服务

35+阅读 · 2020年4月30日

【CMU-TACL2020】低资源跨语言实体链接，Low-resource Crosslingual EntityLinking

专知会员服务

17+阅读 · 2020年3月29日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【论文】多语言神经机器翻译综述（A Comprehensive Survey of Multilingual Neural Machine Translation）

【论文】多语言神经机器翻译综述（A Comprehensive Survey of Multilingual Neural Machine Translation）

专知会员服务

20+阅读 · 2020年1月7日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

热门VIP内容

开通专知VIP会员享更多权益服务

赋能真实世界：基于大语言模型的产业智能体技术、实践与评测综述

军事行动中人工智能系统目标交战的附带损伤评估模型 | 最新文献

【普林斯顿博士论文】面向人本机器人学的安全与学习博弈论融合

美陆军协会（AUSA）2025 年会公布的美国十大武器与防务产品创新

相关资讯

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

上百种预训练中文词向量：Chinese-Word-Vectors

上百种预训练中文词向量：Chinese-Word-Vectors

AINLP

23+阅读 · 2019年2月26日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

【论文推荐】最新五篇命名实体识别（NER）相关论文—对抗学习、语料库、深度多任务学习、先验知识、跨语言语义

专知

37+阅读 · 2018年2月21日

【论文推荐】最新5篇信息抽取（IE）相关论文—开放信息抽取、不完整信息、主动学习、越南语、依存分析

【论文推荐】最新5篇信息抽取（IE）相关论文—开放信息抽取、不完整信息、主动学习、越南语、依存分析

专知

12+阅读 · 2018年2月2日

【论文推荐】最新5篇情感分析相关论文—深度学习情感分析综述、情感分析语料库、情感预测性、上下文和位置感知的因子分解模型、LSTM

【论文推荐】最新5篇情感分析相关论文—深度学习情感分析综述、情感分析语料库、情感预测性、上下文和位置感知的因子分解模型、LSTM

专知

55+阅读 · 2018年1月28日

【推荐】深度学习情感分析综述

【推荐】深度学习情感分析综述

机器学习研究会

58+阅读 · 2018年1月26日

相关论文

Weakly-Supervised Conditional Embedding for Referred Visual Search

Arxiv

0+阅读 · 2023年6月5日

The State of the Art in Creating Visualization Corpora for Automated Chart Analysis

Arxiv

0+阅读 · 2023年6月4日

UCAS-IIE-NLP at SemEval-2023 Task 12: Enhancing Generalization of Multilingual BERT for Low-resource Sentiment Analysis

Arxiv

0+阅读 · 2023年6月1日

Neural Natural Language Processing for Long Texts: A Survey of the State-of-the-Art

Arxiv

0+阅读 · 2023年6月1日

BeamSearchQA: Large Language Models are Strong Zero-Shot QA Solver

Arxiv

0+阅读 · 2023年6月1日

Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili

Arxiv

0+阅读 · 2023年6月1日

Meta Learning for Natural Language Processing: A Survey

Meta Learning for Natural Language Processing: A Survey

Arxiv

14+阅读 · 2022年5月3日

Pretrained Transformers for Text Ranking: BERT and Beyond

Arxiv

28+阅读 · 2020年10月13日

Multilingual Sentiment Analysis: An RNN-Based Framework for Limited Data

Arxiv

12+阅读 · 2018年6月8日

EventKG: A Multilingual Event-Centric Temporal Knowledge Graph

Arxiv

11+阅读 · 2018年4月12日

相关基金

新疆地区维吾尔族、汉族klotho基因单核苷酸多态性与特发性草酸钙肾结石相关性研究

国家自然科学基金

0+阅读 · 2014年12月31日

钬、镨双掺杂氟化镥锂晶体2.9μm中红外激光特性研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于紫外光谱的太阳耀斑能量和动力学过程研究

国家自然科学基金

0+阅读 · 2012年12月31日

CXCR7/SDF-1/ITAC信号调控前列腺癌细胞迁徙、侵袭及增殖的作用机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

复域差分, 差分方程和微分方程的研究

国家自然科学基金

0+阅读 · 2011年12月31日

中文自动口语摘要技术研究

国家自然科学基金

1+阅读 · 2011年12月31日

家庭高等教育投资行为实证研究

国家自然科学基金

0+阅读 · 2011年12月31日

可调谐拍长线扫描的不连续表面三维干涉测量研究

国家自然科学基金

0+阅读 · 2009年12月31日

赖氨酸特异性去甲基酶1对前列腺癌雄激素非依赖性进展的影响及机制

国家自然科学基金

0+阅读 · 2009年12月31日

三波段双极化共用口径SAR天线阵的研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员