Identification of biomarkers from mass spectrometry data using a "common" peak approach

Identification of biomarkers from mass spectrometry data using a "common" peak approach
复制标题

DOI:
10.1186/1471-2105-7-358
复制
发表时间:
2006-07-26
期刊:
影响因子:
3
通讯作者:
Eguchi, Shinto
Eguchi, Shinto
中科院分区:
生物学4区
文献类型:
--
作者:
Fushiki, Tadayoshi;Fujisawa, Hironori;Eguchi, Shinto

文献摘要

被引文献

相似文献

背景:从质谱获得的蛋白质组学数据在早期癌症的检测中引起了极大的兴趣。然而,由于质谱数据是高维的,生物标志物的识别是一个关键问题。分析过程如下:数据预处理、生物标志物的识别以及应用AdaBoost构建分类函数。信息性的“共同”峰由AdaBoost选择。AsymBoost还被检查以平衡假阴性和假阳性。使用卵巢癌dataset.Conclusion:连续协变量和离散协变量可以在本方法中使用的方法的有效性。详细研究了连续协变量和离散协变量结果之间的差异。在这里考虑的例子中,两个协变量都提供了很好的预测,但似乎它们提供了不同类型的信息。我们可以通过整合这两个结果来获得更多关于数据结构的信息。
Background: Proteomic data obtained from mass spectrometry have attracted great interest for the detection of early-stage cancer. However, as mass spectrometry data are high-dimensional, identification of biomarkers is a key problem.Results: This paper proposes the use of "common" peaks in data as biomarkers. Analysis is conducted as follows: data preprocessing, identification of biomarkers, and application of AdaBoost to construct a classification function. Informative "common" peaks are selected by AdaBoost. AsymBoost is also examined to balance false negatives and false positives. The effectiveness of the approach is demonstrated using an ovarian cancer dataset.Conclusion: Continuous covariates and discrete covariates can be used in the present approach. The difference between the result for the continuous covariates and that for the discrete covariates was investigated in detail. In the example considered here, both covariates provide a good prediction, but it seems that they provide different kinds of information. We can obtain more information on the structure of the data by integrating both results.