A data review and re-assessment of ovarian cancer serum proteomic profiling.

A data review and re-assessment of ovarian cancer serum proteomic profiling.
复制标题

DOI:
10.1186/1471-2105-4-24
复制
发表时间:
2003-06-09
期刊:
影响因子:
3
通讯作者:
Zhan M
Zhan M
中科院分区:
生物学4区
文献类型:
--
作者:
Sorace JM;Zhan M

文献摘要

被引文献

相似文献

卵巢癌的早期发现有可能大大降低死亡率。最近,结合先进的数据挖掘算法,使用质谱法来开发患者血清蛋白谱已被报道为实现这一目标的一种有前途的方法。在本报告中,我们分析了从临床蛋白质组学计划数据库网站下载的卵巢数据集8-7-02,使用非参数统计和逐步判别分析来制定诊断患者的规则,并了解数据中的一般模式,可能指导未来的研究。来自癌症和对照组的质谱血清谱显示出许多统计差异。例如,使用Wilcoxon测试比较癌症和对照组之间15,154个质量电荷(M/Z)值的强度,结果检测出3,591个M/Z值,其强度相差10-6或更小的p值。肿瘤与对照间统计差异最大的M/Z值区域出现在M/Z值小于500处。例如,M/Z值为2.7921478和245.53704可用于将癌症与对照组显著区分开。另外三组M/Z值是使用一个训练集开发的,该训练集可以在一个测试集中以100%的灵敏度和特异性区分癌症和对照受试者。基于2.7921478和245.53704的M/Z值区分癌症和对照组的能力表明两组之间存在显著的非生物实验偏倚。这种偏差可能会使使用该数据集寻找可重复诊断价值模式的尝试无效。为了尽量减少错误发现,应该仔细审查使用质谱和数据挖掘算法的结果,并用常规统计方法对其进行基准测试。
The early detection of ovarian cancer has the potential to dramatically reduce mortality. Recently, the use of mass spectrometry to develop profiles of patient serum proteins, combined with advanced data mining algorithms has been reported as a promising method to achieve this goal. In this report, we analyze the Ovarian Dataset 8-7-02 downloaded from the Clinical Proteomics Program Databank website, using nonparametric statistics and stepwise discriminant analysis to develop rules to diagnose patients, as well as to understand general patterns in the data that may guide future research. The mass spectrometry serum profiles derived from cancer and controls exhibited numerous statistical differences. For example, use of the Wilcoxon test in comparing the intensity at each of the 15,154 mass to charge (M/Z) values between the cancer and controls, resulted in the detection of 3,591 M/Z values whose intensities differed by a p-value of 10-6 or less. The region containing the M/Z values of greatest statistical difference between cancer and controls occurred at M/Z values less than 500. For example the M/Z values of 2.7921478 and 245.53704 could be used to significantly separate the cancer from control groups. Three other sets of M/Z values were developed using a training set that could distinguish between cancer and control subjects in a test set with 100% sensitivity and specificity. The ability to discriminate between cancer and control subjects based on the M/Z values of 2.7921478 and 245.53704 reveals the existence of a significant non-biologic experimental bias between these two groups. This bias may invalidate attempts to use this dataset to find patterns of reproducible diagnostic value. To minimize false discovery, results using mass spectrometry and data mining algorithms should be carefully reviewed and benchmarked with routine statistical methods.