MetaboListem and TABoLiSTM: Two Deep Learning Algorithms for Metabolite Named Entity Recognition

MetaboListem and TABoLiSTM: Two Deep Learning Algorithms for Metabolite Named Entity Recognition
复制标题

DOI:
10.1101/2022.02.22.481457
复制
发表时间:
2022-02
期刊:
影响因子:
4.1
通讯作者:
Cheng S. Yeung;Tim Beck;J. Posma
Cheng S. Yeung;Tim Beck;J. Posma
中科院分区:
生物学3区
文献类型:
--
作者:
Cheng S. Yeung;Tim Beck;J. Posma

文献摘要

相似文献

由于相关期刊文献的快速扩展,审查代谢组学文献变得越来越困难。因此需要文本挖掘技术来促进更有效的文献综述。在这里,我们贡献了代谢组学研究全文出版物的标准化语料库,并描述了两种新的代谢物命名实体识别(NER)方法的开发。我们介绍了两种基于双向长短期记忆 (BiLSTM) 网络并结合不同迁移学习技术的代谢物 NER 深度学习方法。我们的第一个模型 (MetaboListem) 遵循使用 GloVe 词嵌入的先前方法。我们的第二个模型利用 BERT 和 BioBERT 进行嵌入,并命名为 TABoLiSTM(Transformer-Affixed BiLSTM)。这些方法在使用基于规则的方法注释的新语料库上进行训练,并在手动注释的代谢组学文章上进行评估。 MetaboListem(F1 得分 0.890、精度 0.892、召回率 0.888)和 TABoLiSTM(BioBERT 版本:F1 得分 0.909、精度 0.926、召回率 0.893)在代谢物 NER 上实现了最先进的性能。创建了包含超过 1,200 个全文开放获取代谢组学出版物和超过 116,000 个带注释的代谢物的语料库。这项工作表明深度学习算法能够准确有效地识别文本中的代谢物名称。所提出的语料库和 NER 算法可用于代谢组学文本挖掘任务,例如信息检索、文档分类和基于文献的发现。可用性 语料库和 NER 算法可免费获取,并附有来自 Github(https://github.com/omicsNLP/MetaboliteNER)的详细说明。
Reviewing the metabolomics literature is becoming increasingly difficult because of the rapid expansion of relevant journal literature. Text-mining technologies are therefore needed to facilitate more efficient literature review. Here we contribute a standardised corpus of full-text publications from metabolomics studies and describe the development of two new metabolite named entity recognition (NER) methods. We introduce two deep learning methods for metabolite NER based on Bidirectional Long Short-Term Memory (BiLSTM) networks incorporating different transfer learning techniques. Our first model (MetaboListem) follows prior methodology using GloVe word embeddings. Our second model exploits BERT and BioBERT for embedding and is named TABoLiSTM (Transformer-Affixed BiLSTM). The methods are trained on a novel corpus annotated using rule-based methods, and evaluated on manually annotated metabolomics articles. MetaboListem (F1 score 0.890, precision 0.892, recall 0.888) and TABoLiSTM (BioBERT version: F1 score 0.909, precision 0.926, recall 0.893) have achieved state-of-the-art performance on metabolite NER. A corpus with >1,200 full-text Open Access metabolomics publications and >116,000 annotated metabolites was created. This work demonstrates that deep learning algorithms are capable of identifying metabolite names accurately and efficiently in text. The proposed corpus and NER algorithms can be used for metabolomics text-mining tasks such as information retrieval, document classification and literature-based discovery. Availability The corpus and NER algorithms are freely available with detailed instructions from Github at https://github.com/omicsNLP/MetaboliteNER.