Computational prediction of human proteins that can be secreted into the bloodstream

Computational prediction of human proteins that can be secreted into the bloodstream
复制标题

DOI:
10.1093/bioinformatics/btn418
复制
发表时间:
2008-10-15
期刊:
影响因子:
5.8
通讯作者:
Xu, Ying
Xu, Ying
中科院分区:
生物学3区
文献类型:
--
作者:
Cui, Juan;Liu, Qi;Xu, Ying

文献摘要

被引文献

相似文献

我们提出了一种新的计算方法,用于预测病变人体组织(如癌症)中高表达和异常表达基因中的哪些蛋白质可以分泌到血液中,这为后续血清蛋白质组学研究提供了可能的标记蛋白。解决这一问题的一个主要挑战是,我们对蛋白质在细胞外分泌后的下游定位的了解非常有限,不足以提供有关血液分泌的有用提示。为了绕过这一困难,我们采取了一种数据挖掘方法,首先通过广泛的文献检索,收集已知由于各种病理条件而被分泌到血液中的人类蛋白质,这些蛋白质是由以前的蛋白质组学研究检测到的,然后问这个问题:这些分泌的蛋白质在物理和化学性质、氨基酸序列和结构特征方面有什么共同之处,可以用来预测它们?我们已经确定了一系列与蛋白质分泌相关的特征,如信号肽、跨膜结构域、糖基化位点、无序区域、二级结构含量、疏水性和极性测量。利用这些特征,我们训练了一个基于支持向量机的分类器来预测血液中的蛋白质分泌。在包含98种人类分泌蛋白和6601种非分泌蛋白的大型测试集上,我们的分类器预测灵敏度为90,预测特异性为98。几个额外的数据集被用来进一步评估我们的分类器的性能。在一组122种蛋白质中,由于各种癌症而在人类血液中发现了异常高的丰度,我们的程序预测其中62种是血液分泌蛋白。将我们的程序应用于微阵列基因表达研究中检测到的胃癌和肺癌组织中异常高表达的基因,我们预测13和31分别为血液分泌,提示它们可以分别作为这两种癌症的潜在生物标志物。我们的研究表明,我们的方法可以为连接基因组和蛋白质组学研究提供非常有用的信息,以发现疾病生物标志物。我们的软件可以访问http://csbl1.bmb.uga.edu/cgi-bin/Secretion/secretion.cgi。
We present a novel computational method for predicting which proteins from highly and abnormally expressed genes in diseased human tissues, such as cancers, can be secreted into the bloodstream, suggesting possible marker proteins for follow-up serum proteomic studies. A main challenging issue in tackling this problem is that our understanding about the downstream localization after proteins are secreted outside the cells is very limited and not sufficient to provide useful hints about secretion to the bloodstream. To bypass this difficulty, we have taken a data mining approach by first collecting, through extensive literature searches, human proteins that are known to be secreted into the bloodstream due to various pathological conditions as detected by previous proteomic studies, and then asking the question: what do these secreted proteins have in common in terms of their physical and chemical properties, amino acid sequence and structural features that can be used to predict them? We have identified a list of features, such as signal peptides, transmembrane domains, glycosylation sites, disordered regions, secondary structural content, hydrophobicity and polarity measures that show relevance to protein secretion. Using these features, we have trained a support vector machine-based classifier to predict protein secretion to the bloodstream. On a large test set containing 98 secretory proteins and 6601 non-secretory proteins of human, our classifier achieved 90 prediction sensitivity and 98 prediction specificity. Several additional datasets are used to further assess the performance of our classifier. On a set of 122 proteins that were found to be of abnormally high abundance in human blood due to various cancers, our program predicted 62 as blood-secreted proteins. By applying our program to abnormally highly expressed genes in gastric cancer and lung cancer tissues detected through microarray gene expression studies, we predicted 13 and 31 as blood secreted, respectively, suggesting that they could serve as potential biomarkers for these two cancers, respectively. Our study demonstrated that our method can provide highly useful information to link genomic and proteomic studies for disease biomarker discovery. Our software can be accessed at http://csbl1.bmb.uga.edu/cgi-bin/Secretion/secretion.cgi.