Deep learning of mutation-gene-drug relations from the literature.

Deep learning of mutation-gene-drug relations from the literature.
复制标题

从文献中深入学习突变 - 毒品的关系。

DOI:
10.1186/s12859-018-2029-1
复制
发表时间:
2018-01-25
期刊:
影响因子:
3
通讯作者:
Kang J
Kang J
中科院分区:
生物学4区
文献类型:
--
作者:
Lee K;Kim B;Choi Y;Kim S;Shin W;Lee S;Park S;Kim S;Tan AC;Kang J

文献摘要

参考文献

被引文献

相似文献

可以预测癌症患者药物疗效的分子生物标志物是推进精准医学的关键组成部分。然而,鉴定这些分子生物标志物仍然是一项艰巨且具有挑战性的任务。下一代患者测序和临床前模型已经越来越多地导致新的基因-突变-药物关系的鉴定,这些结果已经在科学文献中报道和发表。在这里,我们提出了两种新的计算方法,利用所有的PubMed文章作为特定领域的背景知识,以协助从文献中提取和管理基因突变药物的关系。第一种方法使用生物医学实体搜索工具(BEST)评分结果作为训练机器学习分类器的一些特征。第二种方法不仅使用BEST评分结果,还使用深度卷积神经网络模型中的词向量,这些词向量是根据PubMed摘要和Google News文章等大量文档构建和训练的。使用从BEST搜索引擎得分和词向量中获得的特征,我们使用随机森林和深度卷积神经网络等机器学习分类器从文献中提取突变基因和突变药物关系。我们的方法取得了更好的结果相比,国家的最先进的方法。我们在一个简单的机器学习模型中使用了我们提出的特征,并分别获得了0.96和0.82的突变基因和突变药物关系分类的F1分数。我们还使用卷积神经网络、BEST分数和在PubMed或Google News数据上预训练的单词嵌入开发了一个深度学习分类模型。使用深度学习,分类准确率提高,突变-基因和突变-药物关系的F1得分分别为0.96和0.86。我们相信,我们在这项研究中描述的计算方法可以作为一种重要的工具,用于识别预测癌症患者药物反应的分子生物标志物。我们还建立了一个从所有PubMed摘要中提取的突变-基因-药物关系数据库。我们相信,我们的数据库可以成为精准医学研究人员的宝贵资源。本文的在线版本(10.1186/s12859 - 018 - 2029 - 1)包含补充材料,可供授权用户使用。
Molecular biomarkers that can predict drug efficacy in cancer patients are crucial components for the advancement of precision medicine. However, identifying these molecular biomarkers remains a laborious and challenging task. Next-generation sequencing of patients and preclinical models have increasingly led to the identification of novel gene-mutation-drug relations, and these results have been reported and published in the scientific literature. Here, we present two new computational methods that utilize all the PubMed articles as domain specific background knowledge to assist in the extraction and curation of gene-mutation-drug relations from the literature. The first method uses the Biomedical Entity Search Tool (BEST) scoring results as some of the features to train the machine learning classifiers. The second method uses not only the BEST scoring results, but also word vectors in a deep convolutional neural network model that are constructed from and trained on numerous documents such as PubMed abstracts and Google News articles. Using the features obtained from both the BEST search engine scores and word vectors, we extract mutation-gene and mutation-drug relations from the literature using machine learning classifiers such as random forest and deep convolutional neural networks. Our methods achieved better results compared with the state-of-the-art methods. We used our proposed features in a simple machine learning model, and obtained F1-scores of 0.96 and 0.82 for mutation-gene and mutation-drug relation classification, respectively. We also developed a deep learning classification model using convolutional neural networks, BEST scores, and the word embeddings that are pre-trained on PubMed or Google News data. Using deep learning, the classification accuracy improved, and F1-scores of 0.96 and 0.86 were obtained for the mutation-gene and mutation-drug relations, respectively. We believe that our computational methods described in this research could be used as an important tool in identifying molecular biomarkers that predict drug responses in cancer patients. We also built a database of these mutation-gene-drug relations that were extracted from all the PubMed abstracts. We believe that our database can prove to be a valuable resource for precision medicine researchers. The online version of this article (10.1186/s12859-018-2029-1) contains supplementary material, which is available to authorized users.
DOI: 10.1093/bioinformatics/btw511
发表时间: 2016-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lee, Kyubum;Shin, Wonho;Kang, Jaewoo
通讯作者: Kang, Jaewoo
DOI: 10.1186/1758-2946-7-s1-s3
发表时间: 2015
影响因子: 8.6
作者:
Leaman R;Wei CH;Lu Z
通讯作者: Lu Z
DOI: 10.1038/nature11003
发表时间: 2012-03-28
期刊: NATURE
影响因子: 64.8
作者:
Barretina, Jordi;Caponigro, Giordano;Stransky, Nicolas;Venkatesan, Kavitha;Margolin, Adam A.;Kim, Sungjoon;Wilson, Christopher J.;Lehar, Joseph;Kryukov, Gregory V.;Sonkin, Dmitriy;Reddy, Anupama;Liu, Manway;Murray, Lauren;Berger, Michael F.;Monahan, John E.;Morais, Paula;Meltzer, Jodi;Korejwa, Adam;Jane-Valbuena, Judit;Mapa, Felipa A.;Thibault, Joseph;Bric-Furlong, Eva;Raman, Pichai;Shipway, Aaron;Engels, Ingo H.;Cheng, Jill;Yu, Guoying K.;Yu, Jianjun;Aspesi, Peter, Jr.;de Silva, Melanie;Jagtap, Kalpana;Jones, Michael D.;Wang, Li;Hatton, Charles;Palescandolo, Emanuele;Gupta, Supriya;Mahan, Scott;Sougnez, Carrie;Onofrio, Robert C.;Liefeld, Ted;MacConaill, Laura;Winckler, Wendy;Reich, Michael;Li, Nanxin;Mesirov, Jill P.;Gabriel, Stacey B.;Getz, Gad;Ardlie, Kristin;Chan, Vivien;Myer, Vic E.;Weber, Barbara L.;Porter, Jeff;Warmuth, Markus;Finan, Peter;Harris, Jennifer L.;Meyerson, Matthew;Golub, Todd R.;Morrissey, Michael P.;Sellers, William R.;Schlegel, Robert;Garraway, Levi A.
通讯作者: Garraway, Levi A.
DOI: 10.1093/bioinformatics/btv476
发表时间: 2016-01-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Mallory EK;Zhang C;Ré C;Altman RB
通讯作者: Altman RB
DOI: 10.1002/humu.22981
发表时间: 2016-06-01
期刊: HUMAN MUTATION
影响因子: 3.9
作者:
den Dunnen, Johan T.;Dalgleish, Raymond;Taschner, Peter E. M.
通讯作者: Taschner, Peter E. M.