Probabilistic disease classification of expression-dependent proteomic data from mass spectrometry of human serum

Probabilistic disease classification of expression-dependent proteomic data from mass spectrometry of human serum
复制标题

DOI:
10.1089/106652703322756159
复制
发表时间:
2003-01-01
影响因子:
1.7
通讯作者:
Donald, BR
Donald, BR
中科院分区:
生物学4区
文献类型:
--
作者:
Lilien, RH;Farid, H;Donald, BR

文献摘要

被引文献

相似文献

我们开发了一种称为Q5的算法,用于使用质谱法对健康与疾病的全血清样本进行概率分类。该算法采用主成分分析(PCA),其次是线性判别分析(LDA)的全光谱表面增强激光解吸/电离飞行时间(SELDI-TOF)质谱(MS)数据,并证明了四个真实的数据集,从完整的,复杂的SELDI光谱的人血清。Q5是复杂蛋白质混合物的完整质谱分类问题的封闭形式的精确解。Q5采用基于降维线性判别分析的概率分类算法。我们的解决方案是计算效率,它是非迭代的,并使用封闭形式的方程计算最佳线性判别式。最佳判别式的计算和验证的数据集的完整的,复杂的SELDI光谱的人血清。采用每个数据集的不同训练/测试分割的重复实验来验证算法的鲁棒性。概率分类方法取得了优异的性能。我们在三个卵巢癌数据集和一个前列腺癌数据集上实现了97%以上的灵敏度,特异性和阳性预测值。Q5方法优于以前的全光谱复杂样品光谱分类技术,可以提供线索,差异表达的蛋白质和肽的分子身份。
We have developed an algorithm called Q5 for probabilistic classification of healthy versus disease whole serum samples using mass spectrometry. The algorithm employs principal components analysis (PCA) followed by linear discriminant analysis (LDA) on whole spectrum surface-enhanced laser desorption/ionization time of flight (SELDI-TOF) mass spectrometry (MS) data and is demonstrated on four real datasets from complete, complex SELDI spectra of human blood serum. Q5 is a closed-form, exact solution to the problem of classification of complete mass spectra of a complex protein mixture. Q5 employs a probabilistic classification algorithm built upon a dimension-reduced linear discriminant analysis. Our solution is computationally efficient; it is noniterative and computes the optimal linear discriminant using closed-form equations. The optimal discriminant is computed and verified for datasets of complete, complex SELDI spectra of human blood serum. Replicate experiments of different training/testing splits of each dataset are employed to verify robustness of the algorithm. The probabilistic classification method achieves excellent performance. We achieve sensitivity, specificity, and positive predictive values above 97% on three ovarian cancer datasets and one prostate cancer dataset. The Q5 method outperforms previous full-spectrum complex sample spectral classification techniques and can provide clues as to the molecular identities of differentially expressed proteins and peptides.