Computational protein biomarker prediction: a case study for prostate cancer.

Computational protein biomarker prediction: a case study for prostate cancer.
复制标题

DOI:
10.1186/1471-2105-5-26
复制
发表时间:
2004-03-11
期刊:
影响因子:
3
通讯作者:
Wright GL Jr
Wright GL Jr
中科院分区:
生物学4区
文献类型:
--
作者:
Wagner M;Naik DN;Pothen A;Kasukurti S;Devineni RR;Adam BL;Semmes OJ;Wright GL Jr

文献摘要

参考文献

被引文献

相似文献

质谱的最新技术进步对计算数学和统计学提出了挑战,将质谱数据处理成具有临床和生物学意义的预测模型。我们讨论了几种基于分类的方法,使用通过质谱获得的蛋白质谱来寻找蛋白质生物标志物候选物,并评估了它们的统计显著性。我们的总体目标是找出与特定疾病状态具有很高生物学关联可能性的峰值,从而缩小生物标志物候选物的搜索范围。在东弗吉尼亚医学院使用SELDI-TOF质谱法获得的300多名患者的前列腺癌数据集上进行了彻底的交叉验证研究和随机化试验。使用基于两阶段线性支持向量机的过程,我们在四组分类问题上获得了87%的平均分类精度,只有13个峰,与其他方法表现相当。现代特征选择和分类方法是识别候选生物标志物和从蛋白质质谱谱中建立预测模型的相关问题的有力技术。交叉验证和随机化是必不可少的工具,必须仔细执行,以避免结果不公平的偏倚。然而,只有对潜在蛋白质进行生物学验证和鉴定,才能最终确认任何计算预测的实际价值和能力。
Recent technological advances in mass spectrometry pose challenges in computational mathematics and statistics to process the mass spectral data into predictive models with clinical and biological significance. We discuss several classification-based approaches to finding protein biomarker candidates using protein profiles obtained via mass spectrometry, and we assess their statistical significance. Our overall goal is to implicate peaks that have a high likelihood of being biologically linked to a given disease state, and thus to narrow the search for biomarker candidates. Thorough cross-validation studies and randomization tests are performed on a prostate cancer dataset with over 300 patients, obtained at the Eastern Virginia Medical School using SELDI-TOF mass spectrometry. We obtain average classification accuracies of 87% on a four-group classification problem using a two-stage linear SVM-based procedure and just 13 peaks, with other methods performing comparably. Modern feature selection and classification methods are powerful techniques for both the identification of biomarker candidates and the related problem of building predictive models from protein mass spectrometric profiles. Cross-validation and randomization are essential tools that must be performed carefully in order not to bias the results unfairly. However, only a biological validation and identification of the underlying proteins will ultimately confirm the actual value and power of any computational predictions.
DOI: 10.1002/pmic.200300514
发表时间: 2003-09-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Howard, BA;Wang, MZ;Patz, EF
通讯作者: Patz, EF
DOI: 10.1287/opre.13.3.444
发表时间: 1965-01-01
影响因子: 2.7
作者:
MANGASARIAN, OL
通讯作者: MANGASARIAN, OL
DOI: 10.1089/106652703322756159
发表时间: 2003-01-01
影响因子: 1.7
作者:
Lilien, RH;Farid, H;Donald, BR
通讯作者: Donald, BR
DOI: 10.1016/s0140-6736(02)07746-2
发表时间: 2002-02-16
期刊: LANCET
影响因子: 168.9
作者:
Petricoin, EF;Ardekani, AM;Liotta, LA
通讯作者: Liotta, LA
DOI: 10.1186/1471-2105-4-24
发表时间: 2003-06-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Sorace JM;Zhan M
通讯作者: Zhan M