Increasing metadata coverage of SRA BioSample entries using deep learning-based named entity recognition.

Increasing metadata coverage of SRA BioSample entries using deep learning-based named entity recognition.
复制标题

DOI:
10.1093/database/baab021
复制
发表时间:
2021-04-29
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Carter H
Carter H
中科院分区:
其他
文献类型:
--
作者:
Klie A;Tsui BY;Mollah S;Skola D;Dow M;Hsu CN;Carter H

文献摘要

参考文献

被引文献

相似文献

对于大型公共存储库中托管的数据,高质量的元数据注释对于研究的可重复性以及进行快速,强大和可扩展的元分析至关重要。目前,美国国家生物技术信息中心的序列读取档案(SRA)中的大多数测序样本都缺少几个类别的元数据。为了提高这些样本的元数据覆盖率,我们利用来自SRA BioSample的近4400万个属性-值对来训练一个可扩展的递归神经网络,该网络通过命名实体识别(NER)来预测缺失的元数据。该网络首先根据11个元数据类别对短文本短语进行分类,并分别实现了85.2%和0.977的总体准确率和接收器操作特征曲线下面积。然后,我们应用我们的分类器从样本的较长TITLE属性中预测11个元数据类别,评估一组从模型训练中保留的样本的性能。当从标题中提取样本属/种(94.85%),病症/疾病(95.65%)和菌株(82.03%)时,预测准确率很高,其他类别的准确率较低,缺乏预测,突出了BioSample中当前元数据注释的多个问题。这些结果表明了递归神经网络在基于NER的元数据预测中的实用性,以及本文所述模型在增加BioSample中元数据覆盖率的同时最大限度地减少手动管理需求的潜力。 数据库URL:https://github.com/cartercompbio/PredictMEE
High-quality metadata annotations for data hosted in large public repositories are essential for research reproducibility and for conducting fast, powerful and scalable meta-analyses. Currently, a majority of sequencing samples in the National Center for Biotechnology Information’s Sequence Read Archive (SRA) are missing metadata across several categories. In an effort to improve the metadata coverage of these samples, we leveraged almost 44 million attribute–value pairs from SRA BioSample to train a scalable, recurrent neural network that predicts missing metadata via named entity recognition (NER). The network was first trained to classify short text phrases according to 11 metadata categories and achieved an overall accuracy and area under the receiver operating characteristic curve of 85.2% and 0.977, respectively. We then applied our classifier to predict 11 metadata categories from the longer TITLE attribute of samples, evaluating performance on a set of samples withheld from model training. Prediction accuracies were high when extracting sample Genus/Species (94.85%), Condition/Disease (95.65%) and Strain (82.03%) from TITLEs, with lower accuracies and lack of predictions for other categories highlighting multiple issues with the current metadata annotations in BioSample. These results indicate the utility of recurrent neural networks for NER-based metadata prediction and the potential for models such as the one presented here to increase metadata coverage in BioSample while minimizing the need for manual curation. Database URL: https://github.com/cartercompbio/PredictMEE
DOI: 10.1093/nar/gkr937
发表时间: 2012-01
影响因子: 14.9
作者:
Gostev M;Faulconbridge A;Brandizi M;Fernandez-Banet J;Sarkans U;Brazma A;Parkinson H
通讯作者: Parkinson H
DOI: 10.1100/tsw.2009.57
发表时间: 2009-05-29
影响因子: --
作者:
Brazma A
通讯作者: Brazma A
聚类清理:解决生物医学元数据中数据质量问题的方法
DOI: 10.1186/s12859-017-1832-4
发表时间: 2017-09-18
期刊: BMC bioinformatics
影响因子: 3
作者:
Hu W;Zaveri A;Qiu H;Dumontier M
通讯作者: Dumontier M
DOI: 10.1093/nar/30.1.207
发表时间: 2002-01-01
影响因子: 14.9
作者:
Edgar, R;Domrachev, M;Lash, AE
通讯作者: Lash, AE
DOI: 10.1093/bioinformatics/btx334
发表时间: 2017-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bernstein, Matthew N.;Doan, Anhai;Dewey, Colin N.
通讯作者: Dewey, Colin N.