DextMP: deep dive into text for predicting moonlighting proteins.

DextMP: deep dive into text for predicting moonlighting proteins.
复制标题

DOI:
10.1093/bioinformatics/btx231
复制
发表时间:
2017-07-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Kihara D
Kihara D
中科院分区:
其他
文献类型:
--
作者:
Khan IK;Bhuiyan M;Kihara D

文献摘要

参考文献

被引文献

相似文献

兼职蛋白(MP)是一类重要的蛋白质,执行一个以上的独立的细胞功能。近年来,由于发现MP在包括疾病发展在内的各种系统中发挥重要作用,因此越来越受到关注。MP在数据库中的计算函数预测和注释中也有显着的影响。目前,即使在已知蛋白质具有多种不同功能的情况下,MP在生物数据库中也没有被标记。在这项工作中,我们提出了一种名为DextMP的新方法,该方法基于从科学文献和UniProt数据库中提取的文本特征来预测蛋白质是否是MP。DextMP提取蛋白质的三类文本信息:标题、文献摘要和UniProt中的功能描述。应用并比较了三种语言模型:最先进的深度无监督学习算法沿着以及其他两种不同类型的语言模型,词袋中的词频-逆文档频率和主题建模类别中的潜在狄利克雷分配。在已知MP和非MP数据集上的交叉验证结果表明,DextMP成功地预测了MP,准确率超过91%,与现有的MP预测方法相比有显着改进。最后,我们在人类、酵母和非洲爪蟾的三个基因组上使用性能最好的语言模型和基于文本的特征组合运行DextMP,发现大约2.5-35%的蛋白质组是潜在的MP。代码可在http://kiharalab.org/DextMP上获得。
Moonlighting proteins (MPs) are an important class of proteins that perform more than one independent cellular function. MPs are gaining more attention in recent years as they are found to play important roles in various systems including disease developments. MPs also have a significant impact in computational function prediction and annotation in databases. Currently MPs are not labeled as such in biological databases even in cases where multiple distinct functions are known for the proteins. In this work, we propose a novel method named DextMP, which predicts whether a protein is a MP or not based on its textual features extracted from scientific literature and the UniProt database. DextMP extracts three categories of textual information for a protein: titles, abstracts from literature, and function description in UniProt. Three language models were applied and compared: a state-of-the-art deep unsupervised learning algorithm along with two other language models of different types, Term Frequency-Inverse Document Frequency in the bag-of-words and Latent Dirichlet Allocation in the topic modeling category. Cross-validation results on a dataset of known MPs and non-MPs showed that DextMP successfully predicted MPs with over 91% accuracy with significant improvement over existing MP prediction methods. Lastly, we ran DextMP with the best performing language models and text-based feature combinations on three genomes, human, yeast and Xenopus laevis, and found that about 2.5–35% of the proteomes are potential MPs. Code available at http://kiharalab.org/DextMP.
DOI: 10.1186/1471-2105-11-265
发表时间: 2010-05-19
期刊: BMC bioinformatics
影响因子: 3
作者:
Hawkins T;Chitale M;Kihara D
通讯作者: Kihara D
DOI: 10.1186/s13062-014-0030-9
发表时间: 2014-12-11
期刊: Biology direct
影响因子: 5.5
作者:
Khan I;Chen Y;Dong T;Hong X;Takeuchi R;Mori H;Kihara D
通讯作者: Kihara D
DOI: 10.1110/ps.062153506
发表时间: 2006-06-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Hawkins, Troy;Luban, Stanislav;Kihara, Daisuke
通讯作者: Kihara, Daisuke
DOI: 10.1371/journal.pone.0005313
发表时间: 2009
期刊: PloS one
影响因子: 3.7
作者:
Dotan-Cohen D;Letovsky S;Melkman AA;Kasif S
通讯作者: Kasif S
DOI: 10.1142/s0219720007002503
发表时间: 2007-02-01
影响因子: 1
作者:
Hawkins, Troy;Kihara, Daisuke
通讯作者: Kihara, Daisuke