Thousands of samples are needed to generate a robust gene list for predicting outcome in cancer

Thousands of samples are needed to generate a robust gene list for predicting outcome in cancer
复制标题

DOI:
10.1073/pnas.0601231103
复制
发表时间:
2006-04-11
影响因子:
11.1
通讯作者:
Domany, E
Domany, E
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Ein-Dor, L;Zuk, O;Domany, E

文献摘要

被引文献

相似文献

在发现癌症时预测其预后和转移潜力是当前临床研究的主要挑战。最近的许多研究都在寻找优于传统临床参数预测结果的基因表达特征。找到这样一个特征将使许多患者摆脱当前方案下辅助化疗带来的痛苦和毒性,即使他们并不需要这样的治疗。一组可靠的预测基因也将有助于更好地了解转移的生物学机制。一些研究小组已经公布了预测基因的清单,并报告了基于这些基因的良好预测性能。然而,不同群体获得的相同临床类型患者的基因列表差异很大,只有很少的基因是共同的。这种不一致引起了对所报告的预测基因列表的可靠性和稳健性的怀疑,问题的主要来源被证明是用于生成基因列表的样本数量太少。在这里,我们介绍一种以前描述过的数学方法,可能近似正确(PAC)排序,用于评估此类列表的鲁棒性。我们对几个已发表的数据集计算了达到任何期望的可重复性水平所需的样本数量。例如,为了在两个预测基因列表之间实现50%的典型重叠,乳腺癌研究将需要数千名早期发现患者的表达谱。
Predicting at the time of discovery the prognosis and metastatic potential of cancer is a major challenge in current clinical research. Numerous recent studies searched for gene expression signatures that outperform traditionally used clinical parameters in outcome prediction. Finding such a signature will free many patients of the suffering and toxicity associated with adjuvant chemotherapy given to them under current protocols, even though they do not need such treatment. A reliable set of predictive genes also will contribute to a better understanding of the biological mechanism of metastasis. Several groups have published lists of predictive genes and reported good predictive performance based on them. However, the gene lists obtained for the same clinical types of patients by different groups differed widely and had only very few genes in common. This lack of agreement raised doubts about the reliability and robustness of the reported predictive gene lists, and the main source of the problem was shown to be the small number of samples that were used to generate the gene lists. Here, we introduce a previously undescribed mathematical method, probably approximately correct (PAC) sorting, for evaluating the robustness of such lists. We calculate for several published data sets the number of samples that are needed to achieve any desired level of reproducibility. For example, to achieve a typical overlap of 50% between two predictive lists of genes, breast cancer studies would need the expression profiles of several thousand early discovery patients.