Auto-CORPus: A Natural Language Processing Tool for Standardizing and Reusing Biomedical Literature.

Auto-CORPus: A Natural Language Processing Tool for Standardizing and Reusing Biomedical Literature.
复制标题

DOI:
10.3389/fdgth.2022.788124
复制
发表时间:
2022
影响因子:
--
通讯作者:
Posma JM
Posma JM
中科院分区:
其他
文献类型:
--
作者:
Beck T;Shorter T;Hu Y;Li Z;Sun S;Popovici CM;McQuibban NAR;Makraduli F;Yeung CS;Rowlands T;Posma JM

文献摘要

被引文献

相似文献

为了使用机器学习和其他自然语言处理(NLP)算法分析大型语料库,需要对语料库进行标准化。BioC格式是一种社区驱动的简单数据结构,用于共享文本和注释,但BioC格式的生物医学文献访问有限,并且缺乏将在线出版物HTML格式转换为BioC的生物信息学工具。我们提出了Auto-CORPus(研究出版物一致输出的自动管道),这是一种新颖的NLP工具,用于将出版物HTML和表格图像文件标准化和转换为三种方便的机器可解释的输出,以支持生物医学文本分析。首先,Auto-CORPus可以配置为将HTML从各种出版物源转换为BioC。为了标准化异构出版物章节的描述,使用信息本体来注释BioC输出中的每个章节。其次,Auto-CORPus将发布表转换为JSON格式,以便在文本分析系统之间存储、交换和注释表数据。BioC规范不包括用于表示发布表数据的数据结构,因此我们提出了用于共享表内容和元数据的JSON格式。处理全文HTML文件中的内联表和单独HTML文件中的链接表,并将其转换为机器可解释的JSON格式。最后,Auto-CORPus提取出版物文本中声明的缩写,并提供将缩写与完整定义关联的缩写JSON输出。此缩写集合支持文本挖掘任务,例如命名实体识别,方法是包含标准生物本体和字典中不包含的单个出版物特有的缩写。Auto-CORPus软件包可从GitHub免费获得,并附有详细说明:https://github.com/omicsNLP/Auto-CORPus。
To analyse large corpora using machine learning and other Natural Language Processing (NLP) algorithms, the corpora need to be standardized. The BioC format is a community-driven simple data structure for sharing text and annotations, however there is limited access to biomedical literature in BioC format and a lack of bioinformatics tools to convert online publication HTML formats to BioC. We present Auto-CORPus (Automated pipeline for Consistent Outputs from Research Publications), a novel NLP tool for the standardization and conversion of publication HTML and table image files to three convenient machine-interpretable outputs to support biomedical text analytics. Firstly, Auto-CORPus can be configured to convert HTML from various publication sources to BioC. To standardize the description of heterogenous publication sections, the Information Artifact Ontology is used to annotate each section within the BioC output. Secondly, Auto-CORPus transforms publication tables to a JSON format to store, exchange and annotate table data between text analytics systems. The BioC specification does not include a data structure for representing publication table data, so we present a JSON format for sharing table content and metadata. Inline tables within full-text HTML files and linked tables within separate HTML files are processed and converted to machine-interpretable table JSON format. Finally, Auto-CORPus extracts abbreviations declared within publication text and provides an abbreviations JSON output that relates an abbreviation with the full definition. This abbreviation collection supports text mining tasks such as named entity recognition by including abbreviations unique to individual publications that are not contained within standard bio-ontologies and dictionaries. The Auto-CORPus package is freely available with detailed instructions from GitHub at: https://github.com/omicsNLP/Auto-CORPus.