Reference ontology and database annotation of the COVID-19 Open Research Dataset (CORD-19)

Reference ontology and database annotation of the COVID-19 Open Research Dataset (CORD-19)
复制标题

COVID-19 开放研究数据集 (CORD-19) 的参考本体和数据库注释

DOI:
10.1101/2020.10.04.325266
复制
发表时间:
2020
期刊:
bioRxiv
影响因子:
--
通讯作者:
J. Malone
J. Malone
中科院分区:
--
文献类型:
--
作者:
O. Giles;R. Huntley;Anneli Karlsson;J. Lomax;J. Malone

文献摘要

被引文献

相似文献

COVID-19 开放研究数据集 (CORD-19) 于 2020 年 3 月发布,允许机器学习和更广泛的研究社区开发技术来回答有关 COVID-19 的科学问题。该数据集包含大量科学文献,包括超过 100,000 篇全文论文。对训练数据进行注释以标准化生物实体的变异性可以提高下游分析和解释的性能。为了促进和增强 CORD-19 数据在这些应用中的使用,我们于 2020 年 3 月下旬使用命名实体识别工具 TERMite 以及许多大型参考本体和词汇表(包括基因、蛋白质、药物和病毒株领域)执行了全面的注释过程。附加注释已识别并标记了由 62,746 个独特生物医学实体组成的语料库中超过 4500 万个实体。带注释数据的最新更新版本以及旧版本在 GPL-2.0 许可证下公开提供,供社区使用:https://github.com/SciBiteLabs/CORD19
The COVID-19 Open Research Dataset (CORD-19) was released in March 2020 to allow the machine learning and wider research community to develop techniques to answer scientific questions on COVID-19. The dataset consists of a large collection of scientific literature, including over 100,000 full text papers. Annotating training data to normalise variability in biological entities can improve the performance of downstream analysis and interpretation. To facilitate and enhance the use of the CORD-19 data in these applications, in late March 2020 we performed a comprehensive annotation process using named entity recognition tool, TERMite, along with a number of large reference ontologies and vocabularies including domains of genes, proteins, drugs and virus strains. The additional annotation has identified and tagged over 45 million entities within the corpus made up of 62,746 unique biomedical entities. The latest updated version of the annotated data, as well as older versions, is made openly available under GPL-2.0 License for the community to use at: https://github.com/SciBiteLabs/CORD19