生物医学文献中提及的大量软件数据集 (A large dataset of software mentions in the biomedical literature) - 专知论文

会员服务 ·

0

PubMed · 数据集 · entity · 命名实体识别 · 论文 ·

2022 年 9 月 1 日

A large dataset of software mentions in the biomedical literature

翻译：生物医学文献中提及的大量软件数据集

Ana-Maria Istrate,Donghui Li,Dario Taraborelli,Michaela Torkar,Boris Veytsman,Ivana Williams

We describe a new dataset of software mentions in biomedical papers. Plain-text software mentions are extracted with a trained SciBERT model from several sources: the NIH PubMed Central collection and from papers provided by various publishers to the Chan Zuckerberg Initiative. The dataset provides sources, context and metadata, and, for a number of mentions, the disambiguated software entities and links. We extract 1.12 million unique string software mentions from 2.4 million papers in the NIH PMC-OA Commercial subset, 481k unique mentions from the NIH PMC-OA Non-Commercial subset (both gathered in October 2021) and 934k unique mentions from 3 million papers in the Publishers' collection. There is variation in how software is mentioned in papers and extracted by the NER algorithm. We propose a clustering-based disambiguation algorithm to map plain-text software mentions into distinct software entities and apply it on the NIH PubMed Central Commercial collection. Through this methodology, we disambiguate 1.12 million unique strings extracted by the NER model into 97600 unique software entities, covering 78% of all software-paper links. We link 185000 of the mentions to a repository, covering about 55% of all software-paper links. We describe in detail the process of building the datasets, disambiguating and linking the software mentions, as well as opportunities and challenges that come with a dataset of this size. We make all data and code publicly available as a new resource to help assess the impact of software (in particular scientific open source projects) on science.

翻译：我们描述的是生物医学论文中提及的软件的新数据集。平文本软件引用了来自以下几个来源的经过培训的 SciBERT 模型: NIH PubMed Central 集和各出版商向Chan Zuckerberg 倡议提供的论文。数据集提供了源、上下文和元数据, 以及一些隐含的软件实体和链接。我们从NIH PMC-OA 商业子集中的240万篇论文中提取了112万个独特的字符串。 481k 独有的引用来自NIH PMC-OA Non-Commercial子集( 两者均于2021年10月收集) 和934k 独有的引用来自出版商收藏的300万份论文。在纸张中和NER 算法中如何引用软件, 提供了各种源、背景、直线、直线和中央商业收藏。我们用NER 模型所提取的112万个独有的字符串, 包括了所有软体的软件的78%的大小。我们用软件链接, 将所有数据库链接连接到185 。

0

相关内容

PubMed

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

专知会员服务

30+阅读 · 2022年3月8日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

专知会员服务

69+阅读 · 2020年1月2日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

光皮桦OFP基因在次生壁形成中的功能及调控机制

国家自然科学基金

0+阅读 · 2014年12月31日

GTAT4和Myocardin相互作用调控心肌肥厚

国家自然科学基金

0+阅读 · 2014年12月31日

高温诱导尼罗罗非鱼雌鱼性逆转的分子表观机制

国家自然科学基金

0+阅读 · 2014年12月31日

应用代谢组学方法研究重症急性胰腺炎继发MOF的早期预警机制

国家自然科学基金

0+阅读 · 2013年12月31日

晚钠电流触发细胞凋亡致心脏传导疾病的机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

Catestatin蛋白肽段抑制动脉粥样硬化的作用及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

氟-18标记的整合素avb6-半胱氨酸结PET分子探针的研制

国家自然科学基金

0+阅读 · 2012年12月31日

MDSCs在动脉粥样硬化中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

去酰基化ghrelin改善脂肪组织炎症所致胰岛素抵抗的机制- - 调节性T细胞的作用

国家自然科学基金

0+阅读 · 2011年12月31日

探寻与高功能孤独症和Asperger综合征相关的拷贝数变异

国家自然科学基金

0+阅读 · 2009年12月31日

Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature

Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature

Arxiv

0+阅读 · 2022年10月18日

OpenStack and Google Cloud performance comparison in Infrastructure as a Service model

Arxiv

0+阅读 · 2022年10月18日

Object Recognition in Different Lighting Conditions at Various Angles by Deep Learning Method

Arxiv

0+阅读 · 2022年10月18日

Title detection: a novel approach to automatically finding retractions and other editorial notices in the scholarly literature

Arxiv

0+阅读 · 2022年10月18日

ORCA: A Network and Architecture Co-design for Offloading us-scale Datacenter Applications

Arxiv

0+阅读 · 2022年10月17日

Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World

Arxiv

0+阅读 · 2022年10月16日

On the Identifiability of Nonlinear ICA: Sparsity and Beyond

Arxiv

0+阅读 · 2022年10月16日

New Secure Sparse Inner Product with Applications to Machine Learning

Arxiv

0+阅读 · 2022年10月16日

The Promises of Parallel Outcomes

Arxiv

0+阅读 · 2022年10月14日

Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods

Arxiv

0+阅读 · 2022年10月13日

VIP会员

文章信息

相关主题

命名实体识别

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

专知会员服务

30+阅读 · 2022年3月8日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

专知会员服务

69+阅读 · 2020年1月2日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《物联网（IoT）中的无人机通信高效控制》135页

《在GNSS信号降级环境中利用共识实现无人机集群稳健协调》

中程单向攻击无人机的战略意义：俄乌战争启示

《面向无人机集群的避障动态传感器覆盖算法》最新38页

相关资讯

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature

Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature

Arxiv

0+阅读 · 2022年10月18日

OpenStack and Google Cloud performance comparison in Infrastructure as a Service model

Arxiv

0+阅读 · 2022年10月18日

Object Recognition in Different Lighting Conditions at Various Angles by Deep Learning Method

Arxiv

0+阅读 · 2022年10月18日

Title detection: a novel approach to automatically finding retractions and other editorial notices in the scholarly literature

Arxiv

0+阅读 · 2022年10月18日

ORCA: A Network and Architecture Co-design for Offloading us-scale Datacenter Applications

Arxiv

0+阅读 · 2022年10月17日

Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World

Arxiv

0+阅读 · 2022年10月16日

On the Identifiability of Nonlinear ICA: Sparsity and Beyond

Arxiv

0+阅读 · 2022年10月16日

New Secure Sparse Inner Product with Applications to Machine Learning

Arxiv

0+阅读 · 2022年10月16日

The Promises of Parallel Outcomes

Arxiv

0+阅读 · 2022年10月14日

Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods

Arxiv

0+阅读 · 2022年10月13日

相关基金

光皮桦OFP基因在次生壁形成中的功能及调控机制

国家自然科学基金

0+阅读 · 2014年12月31日

GTAT4和Myocardin相互作用调控心肌肥厚

国家自然科学基金

0+阅读 · 2014年12月31日

高温诱导尼罗罗非鱼雌鱼性逆转的分子表观机制

国家自然科学基金

0+阅读 · 2014年12月31日

应用代谢组学方法研究重症急性胰腺炎继发MOF的早期预警机制

国家自然科学基金

0+阅读 · 2013年12月31日

晚钠电流触发细胞凋亡致心脏传导疾病的机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

Catestatin蛋白肽段抑制动脉粥样硬化的作用及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

氟-18标记的整合素avb6-半胱氨酸结PET分子探针的研制

国家自然科学基金

0+阅读 · 2012年12月31日

MDSCs在动脉粥样硬化中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

去酰基化ghrelin改善脂肪组织炎症所致胰岛素抵抗的机制- - 调节性T细胞的作用

国家自然科学基金

0+阅读 · 2011年12月31日

探寻与高功能孤独症和Asperger综合征相关的拷贝数变异

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员