Text Mining Genotype-Phenotype Relationships from Biomedical Literature for Database Curation and Precision Medicine.

Text Mining Genotype-Phenotype Relationships from Biomedical Literature for Database Curation and Precision Medicine.
复制标题

文本挖掘基因型 - 表型关系从生物医学文献用于数据库策展和精确医学。

DOI:
10.1371/journal.pcbi.1005017
复制
发表时间:
2016-11
影响因子:
4.3
通讯作者:
Lu Z
Lu Z
中科院分区:
生物学2区
文献类型:
--
作者:
Singhal A;Simmons M;Lu Z

文献摘要

参考文献

被引文献

相似文献

精准医疗的实践最终需要基因和突变数据库供医疗保健提供者参考,以了解每个患者基因组成的临床意义。虽然最高质量的数据库需要手动管理,但文本挖掘工具可以促进管理过程,提高准确性,覆盖率和生产力。然而,到目前为止,还没有可用的文本挖掘工具,提供高精度的性能,从生物医学文献中提取这样的三联体。在本文中,我们提出了一个高性能的机器学习方法,自动提取疾病基因变异三联体的生物医学文献。我们的方法是独特的,因为我们不仅从本地文本内容中识别与每个突变相关的基因和蛋白质产物,而且还从全球范围内(来自互联网和PubMed中的所有文献)。我们的方法还采用了一种新的基于文本挖掘的机器学习方法,结合了蛋白质序列验证和疾病关联。我们从PubMed的所有摘要中提取疾病基因变异三联体,这些摘要与10种重要疾病(乳腺癌、前列腺癌、胰腺癌、肺癌、急性髓性白血病、阿尔茨海默病、血色病、年龄相关性黄斑变性(AMD)、糖尿病和囊性纤维化)有关。然后,我们以两种方式评估我们的方法:(1)使用基准数据集与最先进的方法进行直接比较;(2)验证研究,将我们的方法的结果与流行的人类策划数据库(UniProt)中的条目进行比较。在基准比较中,我们的完整方法在F1测量中实现了28%的改进(从0.62到0.79),超过了最先进的结果。对于UniProt知识库(KB)的验证研究,我们对结果和错误进行了全面分析。在所有疾病中,我们的方法返回了272个与UniProt中的条目重叠的三联体(疾病基因变体)和5,384个在UniProt中没有重叠的三联体。重叠的三胞胎和分层样本的非重叠三胞胎的分析显示,93%和80%的准确性为各自的类别(累积准确性,77%)。我们的结论是,我们的过程代表了一个重要的和广泛适用的改进,最先进的疾病基因变异关系的治疗。为了提供个性化的医疗保健,重要的是要了解患者的基因组变异和这些变异在保护或诱发患者疾病的影响。有几个项目旨在通过使用来自临床试验和生物医学文献的数据在有组织的数据库中手动管理这种基因型-表型关系来提供这种信息。然而,生物医学文献的数量呈指数级增长,人工管理者发现文本中“隐藏”的基因型-表型关系的能力有限,这导致了数据库更新的延迟。其结果是在利用目前可用于开发个性化医疗保健解决方案的有价值信息方面存在瓶颈。过去,一些计算技术试图通过使用文本挖掘技术从生物医学文献中自动挖掘基因型-表型信息来加速策展工作。然而,这样的计算方法还没有能够达到足够的精度水平,使它们吸引实际使用。在这项工作中,我们提出了一个高度准确的基于机器学习的文本挖掘方法,从生物医学文献中挖掘完整的基因型-表型关系。我们测试了这种方法对10种著名疾病的性能,并证明了我们的方法的有效性及其潜在的实用价值。我们目前正在努力为所有PubMed数据生成基因型-表型关系,目标是开发一个生命科学中所有已知疾病的详尽数据库。我们相信,这项工作将为使用基因组数据实施个性化医疗保健提供非常重要和必要的支持。
The practice of precision medicine will ultimately require databases of genes and mutations for healthcare providers to reference in order to understand the clinical implications of each patient’s genetic makeup. Although the highest quality databases require manual curation, text mining tools can facilitate the curation process, increasing accuracy, coverage, and productivity. However, to date there are no available text mining tools that offer high-accuracy performance for extracting such triplets from biomedical literature. In this paper we propose a high-performance machine learning approach to automate the extraction of disease-gene-variant triplets from biomedical literature. Our approach is unique because we identify the genes and protein products associated with each mutation from not just the local text content, but from a global context as well (from the Internet and from all literature in PubMed). Our approach also incorporates protein sequence validation and disease association using a novel text-mining-based machine learning approach. We extract disease-gene-variant triplets from all abstracts in PubMed related to a set of ten important diseases (breast cancer, prostate cancer, pancreatic cancer, lung cancer, acute myeloid leukemia, Alzheimer’s disease, hemochromatosis, age-related macular degeneration (AMD), diabetes mellitus, and cystic fibrosis). We then evaluate our approach in two ways: (1) a direct comparison with the state of the art using benchmark datasets; (2) a validation study comparing the results of our approach with entries in a popular human-curated database (UniProt) for each of the previously mentioned diseases. In the benchmark comparison, our full approach achieves a 28% improvement in F1-measure (from 0.62 to 0.79) over the state-of-the-art results. For the validation study with UniProt Knowledgebase (KB), we present a thorough analysis of the results and errors. Across all diseases, our approach returned 272 triplets (disease-gene-variant) that overlapped with entries in UniProt and 5,384 triplets without overlap in UniProt. Analysis of the overlapping triplets and of a stratified sample of the non-overlapping triplets revealed accuracies of 93% and 80% for the respective categories (cumulative accuracy, 77%). We conclude that our process represents an important and broadly applicable improvement to the state of the art for curation of disease-gene-variant relationships. To provide personalized health care it is important to understand patients’ genomic variations and the effect these variants have in protecting or predisposing patients to disease. Several projects aim at providing this information by manually curating such genotype-phenotype relationships in organized databases using data from clinical trials and biomedical literature. However, the exponentially increasing size of biomedical literature and the limited ability of manual curators to discover the genotype-phenotype relationships “hidden” in text has led to delays in keeping such databases updated with the current findings. The result is a bottleneck in leveraging valuable information that is currently available to develop personalized health care solutions. In the past, a few computational techniques have attempted to speed up the curation efforts by using text mining techniques to automatically mine genotype-phenotype information from biomedical literature. However, such computational approaches have not been able to achieve accuracy levels sufficient to make them appealing for practical use. In this work, we present a highly accurate machine-learning-based text mining approach for mining complete genotype-phenotype relationships from biomedical literature. We test the performance of this approach on ten well-known diseases and demonstrate the validity of our approach and its potential utility for practical purposes. We are currently working towards generating genotype-phenotype relationships for all PubMed data with the goal of developing an exhaustive database of all the known diseases in life science. We believe that this work will provide very important and needed support for implementation of personalized health care using genomic data.
DOI: 10.1186/1471-2164-11-s4-s24
发表时间: 2010-12-02
期刊: BMC genomics
影响因子: 4.4
作者:
Laurila JB;Naderi N;Witte R;Riazanov A;Kouznetsov A;Baker CJ
通讯作者: Baker CJ
DOI: 10.1002/humu.22594
发表时间: 2014-08
期刊: HUMAN MUTATION
影响因子: 3.9
作者:
Famiglietti, Maria Livia;Estreicher, Anne;Gos, Arnaud;Bolleman, Jerven;Gehant, Sebastien;Breuza, Lionel;Bridge, Alan;Poux, Sylvain;Redaschi, Nicole;Bougueleret, Lydie;Xenarios, Ioannis
通讯作者: Xenarios, Ioannis
DOI: 10.1093/bioinformatics/btl421
发表时间: 2006-10-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bonis, Julio;Furlong, Laura Ines;Sanz, Ferran
通讯作者: Sanz, Ferran
DOI: 10.1002/humu.21317
发表时间: 2010-09-01
期刊: HUMAN MUTATION
影响因子: 3.9
作者:
Kuipers, Remko;van den Bergh, Tom;Schaap, Peter J.
通讯作者: Schaap, Peter J.
DOI: 10.1093/nar/gku989
发表时间: 2015-01
影响因子: 14.9
作者:
UniProt Consortium
通讯作者: UniProt Consortium