Accelerated variant curation from scientific literature using biomedical text mining.

Accelerated variant curation from scientific literature using biomedical text mining.
复制标题

DOI:
10.17912/micropub.biology.000578
复制
发表时间:
2022
影响因子:
--
通讯作者:
Howe, Kevin
Howe, Kevin
中科院分区:
其他
文献类型:
--
作者:
Mallick, Rishab;Arnaboldi, Valerio;Davis, Paul;Diamantakis, Stavros;Zarowiecki, Magdalena;Howe, Kevin

文献摘要

相似文献

生物数据库通过生物定位收集和标准化数据。尽管主要的模式生物数据库已经采用了一些自动化的管理方法,但大部分生物管理仍然是手动进行的。为了加快变体基因组位置的提取,我们开发了一种混合方法,该方法结合了正则表达式,基于BERT(来自变压器的双向编码器表示)的命名实体识别和词袋,以从C中提取变体基因组位置。为WormBase提供的优雅论文。我们的模型对100篇论文中提取的文本进行了基因突变匹配测试,准确率为82.59%,甚至恢复了一些在人工策展过程中没有发现的数据。代码:https://github.com/WormBase/genomic-info-from-papers
Biological databases collect and standardize data through biocuration. Even though major model organism databases have adopted some automation of curation methods, a large portion of biocuration is still performed manually. To speed up the extraction of the genomic positions of variants, we have developed a hybrid approach that combines regular expressions, Named Entity Recognition based on BERT (Bidirectional Encoder Representations from Transformers) and bag-of-words to extract variant genomic locations from C. elegans papers for WormBase. Our model has a precision of 82.59% for the gene-mutation matches tested on extracted text from 100 papers, and even recovers some data not discovered during manual curation. Code at: https://github.com/WormBase/genomic-info-from-papers