Text Mining Genotype-Phenotype Relationships from Biomedical Literature for Database Curation and Precision Medicine.
Text Mining Genotype-Phenotype Relationships from Biomedical Literature for Database Curation and Precision Medicine.
复制标题
文本挖掘基因型 - 表型关系从生物医学文献用于数据库策展和精确医学。
DOI:
10.1371/journal.pcbi.1005017
复制
发表时间:
2016-11
影响因子:
4.3
通讯作者:
Lu Z
中科院分区:
文献类型:
--
作者:
Singhal A;Simmons M;Lu Z
The practice of precision medicine will ultimately require databases of genes and mutations for healthcare providers to reference in order to understand the clinical implications of each patient’s genetic makeup. Although the highest quality databases require manual curation, text mining tools can facilitate the curation process, increasing accuracy, coverage, and productivity. However, to date there are no available text mining tools that offer high-accuracy performance for extracting such triplets from biomedical literature. In this paper we propose a high-performance machine learning approach to automate the extraction of disease-gene-variant triplets from biomedical literature. Our approach is unique because we identify the genes and protein products associated with each mutation from not just the local text content, but from a global context as well (from the Internet and from all literature in PubMed). Our approach also incorporates protein sequence validation and disease association using a novel text-mining-based machine learning approach. We extract disease-gene-variant triplets from all abstracts in PubMed related to a set of ten important diseases (breast cancer, prostate cancer, pancreatic cancer, lung cancer, acute myeloid leukemia, Alzheimer’s disease, hemochromatosis, age-related macular degeneration (AMD), diabetes mellitus, and cystic fibrosis). We then evaluate our approach in two ways: (1) a direct comparison with the state of the art using benchmark datasets; (2) a validation study comparing the results of our approach with entries in a popular human-curated database (UniProt) for each of the previously mentioned diseases. In the benchmark comparison, our full approach achieves a 28% improvement in F1-measure (from 0.62 to 0.79) over the state-of-the-art results. For the validation study with UniProt Knowledgebase (KB), we present a thorough analysis of the results and errors. Across all diseases, our approach returned 272 triplets (disease-gene-variant) that overlapped with entries in UniProt and 5,384 triplets without overlap in UniProt. Analysis of the overlapping triplets and of a stratified sample of the non-overlapping triplets revealed accuracies of 93% and 80% for the respective categories (cumulative accuracy, 77%). We conclude that our process represents an important and broadly applicable improvement to the state of the art for curation of disease-gene-variant relationships. To provide personalized health care it is important to understand patients’ genomic variations and the effect these variants have in protecting or predisposing patients to disease. Several projects aim at providing this information by manually curating such genotype-phenotype relationships in organized databases using data from clinical trials and biomedical literature. However, the exponentially increasing size of biomedical literature and the limited ability of manual curators to discover the genotype-phenotype relationships “hidden” in text has led to delays in keeping such databases updated with the current findings. The result is a bottleneck in leveraging valuable information that is currently available to develop personalized health care solutions. In the past, a few computational techniques have attempted to speed up the curation efforts by using text mining techniques to automatically mine genotype-phenotype information from biomedical literature. However, such computational approaches have not been able to achieve accuracy levels sufficient to make them appealing for practical use. In this work, we present a highly accurate machine-learning-based text mining approach for mining complete genotype-phenotype relationships from biomedical literature. We test the performance of this approach on ten well-known diseases and demonstrate the validity of our approach and its potential utility for practical purposes. We are currently working towards generating genotype-phenotype relationships for all PubMed data with the goal of developing an exhaustive database of all the known diseases in life science. We believe that this work will provide very important and needed support for implementation of personalized health care using genomic data.
登录
查看更多内容
影响因子:
4.4
作者:
Laurila JB;Naderi N;Witte R;Riazanov A;Kouznetsov A;Baker CJ
通讯作者:
Baker CJ
影响因子:
3.9
作者:
Famiglietti, Maria Livia;Estreicher, Anne;Gos, Arnaud;Bolleman, Jerven;Gehant, Sebastien;Breuza, Lionel;Bridge, Alan;Poux, Sylvain;Redaschi, Nicole;Bougueleret, Lydie;Xenarios, Ioannis
通讯作者:
Xenarios, Ioannis
影响因子:
5.8
作者:
Bonis, Julio;Furlong, Laura Ines;Sanz, Ferran
通讯作者:
Sanz, Ferran
影响因子:
3.9
作者:
Kuipers, Remko;van den Bergh, Tom;Schaap, Peter J.
通讯作者:
Schaap, Peter J.
影响因子:
14.9
作者:
UniProt Consortium
通讯作者:
UniProt Consortium