埃斯科尔比乌斯:西班牙大规模爬行公司 (esCorpius: A Massive Spanish Crawling Corpus) - 专知论文

会员服务 ·

0

WEB · 讲稿 · Extensibility · Processing（编程语言） · URL ·

2022 年 7 月 1 日

esCorpius: A Massive Spanish Crawling Corpus

翻译：埃斯科尔比乌斯:西班牙大规模爬行公司

Asier Gutiérrez-Fandiño,David Pérez-Fernández,Jordi Armengol-Estapé,David Griol,Zoraida Callejas

from arxiv, esCorpius is available on https://huggingface.co/datasets/LHF/escorpius

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in languages other than English. Recently, several initiatives have presented multilingual datasets obtained from automatic web crawling. However, the results in Spanish present important shortcomings, as they are either too small in comparison with other languages, or present a low quality derived from sub-optimal cleaning and deduplication. In this paper, we introduce esCorpius, a Spanish crawling corpus obtained from near 1 Pb of Common Crawl data. It is the most extensive corpus in Spanish with this level of quality in the extraction, purification and deduplication of web textual content. Our data curation process involves a novel highly parallel cleaning pipeline and encompasses a series of deduplication mechanisms that together ensure the integrity of both document and paragraph boundaries. Additionally, we maintain both the source web page URL and the WARC shard origin URL in order to complain with EU regulations. esCorpius has been released under CC BY-NC-ND 4.0 license and is available on HuggingFace.

翻译：近年来,以变压器为基础的模型导致自然语言处理的语言建模取得显著进展,然而,这些模型需要大量的数据(预先)培训,而且缺乏英文以外的语言的Corpora。最近,一些举措提供了自动上网获取的多语种数据集;然而,西班牙文的结果表明存在重大缺陷,因为它们与其他语言相比太小,或来自亚优清洁和淡化的低质量。在本文中,我们引入了Es Corpius,这是一个西班牙爬行体,取自共同Crawl数据近1磅的西班牙Corpius。Es Corpius是根据CC-NC-ND4.0许可证发放的,在网络文本内容的提取、净化和复制方面质量如此之高。我们的数据整理过程涉及一个全新的高度平行的清理管道,包括一系列的解析机制,共同确保文件和段落边界的完整性。此外,我们维护源网页URL和WAC hard URL,以便向欧盟条例投诉。Esorbisius已经根据CC-NC-ND4.0许可证在Hugh Fasing Fasing上发布。

0

相关内容

WEB

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Call for Nominations: 2022 Multimedia Prize Paper Award

Call for Nominations: 2022 Multimedia Prize Paper Award

CCF多媒体专委会

0+阅读 · 2022年2月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

基于PyTorch/TorchText的自然语言处理库

基于PyTorch/TorchText的自然语言处理库

专知

28+阅读 · 2019年4月22日

ClC-3氯通道蛋白在肾上腺素能受体介导心肌肥厚中的功能研究

国家自然科学基金

0+阅读 · 2015年12月31日

细菌角蛋白酶KerF降解角蛋白过程与分子机制

国家自然科学基金

0+阅读 · 2015年12月31日

基于电子束辐照技术降解Musalais中氨基甲酸乙酯的安全性控制机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

ADS中检测快中子束的GEM探测器的研制

国家自然科学基金

0+阅读 · 2013年12月31日

基于MEMS技术的光栅结合可动Fabry-Perot腔微型光谱仪研究

国家自然科学基金

0+阅读 · 2013年12月31日

拟Frobenius-Lusztig核

国家自然科学基金

0+阅读 · 2012年12月31日

联合CORS网络的高分辨率PSI大气延迟定标方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

生理和缺血再灌注状态下的冠脉内皮功能 - - 内皮离子通道间信号关联的研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于Decorin基因甲基化调控的非小细胞肺癌转移的分子机制

国家自然科学基金

0+阅读 · 2011年12月31日

MAWD/MAWBP复合体调节TGF-beta通路的机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

Arxiv

0+阅读 · 2022年8月22日

Representing Knowledge by Spans: A Knowledge-Enhanced Model for Information Extraction

Arxiv

0+阅读 · 2022年8月20日

Beyond Text Generation: Supporting Writers with Continuous Automatic Text Summaries

Arxiv

0+阅读 · 2022年8月19日

Ultron: An Ultimate Retriever on Corpus with a Model-based Indexer

Arxiv

0+阅读 · 2022年8月19日

Quality issues in Machine Learning Software Systems

Arxiv

0+阅读 · 2022年8月18日

Adaptive Re-Ranking with a Corpus Graph

Arxiv

0+阅读 · 2022年8月18日

Low Emission Building Control with Zero-Shot Reinforcement Learning

Arxiv

0+阅读 · 2022年8月18日

MulZDG: Multilingual Code-Switching Framework for Zero-shot Dialogue Generation

Arxiv

0+阅读 · 2022年8月18日

Pre-training Methods in Information Retrieval

Arxiv

16+阅读 · 2021年11月27日

Pre-Training with Whole Word Masking for Chinese BERT

Arxiv

11+阅读 · 2019年6月19日

VIP会员

文章信息

相关主题

Processing（编程语言）

相关VIP内容

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

面向性能、成本效益、云边隐私与可信性的大小语言模型协作综述

乌克兰太空研究（2022-2024年） | 176页

【CMU博士论文】大型语言模型的隐性特性

国防领域人工智能走向何方？

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Call for Nominations: 2022 Multimedia Prize Paper Award

Call for Nominations: 2022 Multimedia Prize Paper Award

CCF多媒体专委会

0+阅读 · 2022年2月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

基于PyTorch/TorchText的自然语言处理库

基于PyTorch/TorchText的自然语言处理库

专知

28+阅读 · 2019年4月22日

相关论文

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

Arxiv

0+阅读 · 2022年8月22日

Representing Knowledge by Spans: A Knowledge-Enhanced Model for Information Extraction

Arxiv

0+阅读 · 2022年8月20日

Beyond Text Generation: Supporting Writers with Continuous Automatic Text Summaries

Arxiv

0+阅读 · 2022年8月19日

Ultron: An Ultimate Retriever on Corpus with a Model-based Indexer

Arxiv

0+阅读 · 2022年8月19日

Quality issues in Machine Learning Software Systems

Arxiv

0+阅读 · 2022年8月18日

Adaptive Re-Ranking with a Corpus Graph

Arxiv

0+阅读 · 2022年8月18日

Low Emission Building Control with Zero-Shot Reinforcement Learning

Arxiv

0+阅读 · 2022年8月18日

MulZDG: Multilingual Code-Switching Framework for Zero-shot Dialogue Generation

Arxiv

0+阅读 · 2022年8月18日

Pre-training Methods in Information Retrieval

Arxiv

16+阅读 · 2021年11月27日

Pre-Training with Whole Word Masking for Chinese BERT

Arxiv

11+阅读 · 2019年6月19日

相关基金

ClC-3氯通道蛋白在肾上腺素能受体介导心肌肥厚中的功能研究

国家自然科学基金

0+阅读 · 2015年12月31日

细菌角蛋白酶KerF降解角蛋白过程与分子机制

国家自然科学基金

0+阅读 · 2015年12月31日

基于电子束辐照技术降解Musalais中氨基甲酸乙酯的安全性控制机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

ADS中检测快中子束的GEM探测器的研制

国家自然科学基金

0+阅读 · 2013年12月31日

基于MEMS技术的光栅结合可动Fabry-Perot腔微型光谱仪研究

国家自然科学基金

0+阅读 · 2013年12月31日

拟Frobenius-Lusztig核

国家自然科学基金

0+阅读 · 2012年12月31日

联合CORS网络的高分辨率PSI大气延迟定标方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

生理和缺血再灌注状态下的冠脉内皮功能 - - 内皮离子通道间信号关联的研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于Decorin基因甲基化调控的非小细胞肺癌转移的分子机制

国家自然科学基金

0+阅读 · 2011年12月31日

MAWD/MAWBP复合体调节TGF-beta通路的机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员