Selecting differentially expressed genes from microarray experiments

Selecting differentially expressed genes from microarray experiments
复制标题

DOI:
10.1111/1541-0420.00016
复制
发表时间:
2003-03-01
期刊:
影响因子:
1.9
通讯作者:
Schummer, M
Schummer, M
中科院分区:
数学3区
文献类型:
--
作者:
Pepe, MS;Longton, G;Schummer, M

文献摘要

被引文献

相似文献

高通量技术,如基因表达阵列和蛋白质质谱,允许人们同时评估数千种可能区分不同组织类型的潜在生物标志物。这里特别感兴趣的是区分癌组织和正常器官组织。我们考虑用统计学方法对基因(或蛋白质)在组织间的差异表达进行排序。各种统计措施被认为是,我们认为,两个措施相关的接收器工作特征曲线是特别适合于此目的。我们还建议量化基因排序中的抽样变异性,并建议使用“选择概率函数”,即每个基因排序的概率分布。这是通过bootstrap估计的。一个真实的数据集,来自23个正常和30个卵巢癌组织的基因表达阵列,进行了分析。模拟研究也被用来评估不同的统计基因排名措施和我们的量化抽样变异的相对性能。我们的方法自然会导致样本量的计算程序,适合探索性研究,寻求确定差异表达的基因。
High throughput technologies, such as gene expression arrays and protein mass spectrometry, allow one to simultaneously evaluate thousands of potential biomarkers that could distinguish different tissue types. Of particular interest here is distinguishing between cancerous and normal organ tissues. We consider statistical methods to rank genes (or proteins) in regards to differential expression between tissues. Various statistical measures are considered, and we argue that two measures related to the Receiver Operating Characteristic Curve are particularly suitable for this purpose. We also propose that sampling variability in the gene rankings be quantified, and suggest using the "selection probability function," the probability distribution of rankings for each gene. This is estimated via the bootstrap. A real dataset, derived from gene expression arrays of 23 normal and 30 ovarian cancer tissues, is analyzed. Simulation studies are also used to assess the relative performance of different statistical gene ranking measures and our quantification of sampling variability. Our approach leads naturally to a procedure for sample-size calculations, appropriate for exploratory studies that seek to identify differentially expressed genes.