Preferred analysis methods for Affymetrix GeneChips revealed by a wholly defined control dataset.

Preferred analysis methods for Affymetrix GeneChips revealed by a wholly defined control dataset.
复制标题

DOI:
10.1186/gb-2005-6-2-r16
复制
发表时间:
2005
期刊:
影响因子:
12.3
通讯作者:
Halfon MS
Halfon MS
中科院分区:
生物学1区
文献类型:
--
作者:
Choe SE;Boutros M;Michelson AM;Church GM;Halfon MS

文献摘要

参考文献

被引文献

相似文献

描述了 Affymetrix GeneChips 的“spike-in”实验,该实验提供了 3,860 个 RNA 种类的定义数据集。提出了分析方法的“最佳途径”组合,允许在达到 10% 的错误发现率之前检测到大约 70% 的真阳性。随着越来越多的方法被开发来分析 RNA 分析数据,使用对照数据集评估其性能变得越来越重要。我们提出了 Affymetrix GeneChips 的“spike-in”实验,该实验提供了 3,860 个 RNA 物种的定义数据集,我们用它来评估识别差异表达基因的分析选项。该实验设计包含两个新颖的特点。首先,为了获得假阳性和假阴性率的准确估计,在每个感兴趣的倍数变化水平(范围从 1.2 到 4 倍)加入 100-200 个 RNA。其次,不是使用未表征的背景 RNA 样本,而是使用一组 2,551 个 RNA 物种作为恒定 (1x) 集,使我们能够知道任何给定的探针集是否真正存在或不存在。对该数据集应用大量分析方法揭示了它们识别差异表达基因的能力的明显差异。当选择以下选项时,假阴性和假阳性率可降至最低:从 PM 探针强度中减去非特异性信号;在探针组水平上执行强度依赖性标准化;并在测试统计中纳入信号强度相关的标准偏差。提出了分析方法的最佳路径组合,允许在达到 10% 的错误发现率之前检测到大约 70% 的真阳性。我们强调了需要改进的领域,包括更好地估计错误发现率和降低假阴性率。
A 'spike-in' experiment for Affymetrix GeneChips is described that provides a defined dataset of 3,860 RNA species. A 'best route' combination of analysis methods is presented which allows detection of approximately 70% of true positives before reaching a 10% false discovery rate. As more methods are developed to analyze RNA-profiling data, assessing their performance using control datasets becomes increasingly important. We present a 'spike-in' experiment for Affymetrix GeneChips that provides a defined dataset of 3,860 RNA species, which we use to evaluate analysis options for identifying differentially expressed genes. The experimental design incorporates two novel features. First, to obtain accurate estimates of false-positive and false-negative rates, 100-200 RNAs are spiked in at each fold-change level of interest, ranging from 1.2 to 4-fold. Second, instead of using an uncharacterized background RNA sample, a set of 2,551 RNA species is used as the constant (1x) set, allowing us to know whether any given probe set is truly present or absent. Application of a large number of analysis methods to this dataset reveals clear variation in their ability to identify differentially expressed genes. False-negative and false-positive rates are minimized when the following options are chosen: subtracting nonspecific signal from the PM probe intensities; performing an intensity-dependent normalization at the probe set level; and incorporating a signal intensity-dependent standard deviation in the test statistic. A best-route combination of analysis methods is presented that allows detection of approximately 70% of true positives before reaching a 10% false-discovery rate. We highlight areas in need of improvement, including better estimate of false-discovery rates and decreased false-negative rates.
DOI: 10.1073/pnas.011404098
发表时间: 2001-01-02
影响因子: 11.1
作者:
Li, C;Wong, WH
通讯作者: Wong, WH
DOI: 10.1073/pnas.091062498
发表时间: 2001-04-24
影响因子: 11.1
作者:
Tusher, VG;Tibshirani, R;Chu, G
通讯作者: Chu, G
DOI: 10.1002/jcb.10073
发表时间: 2001-01-01
影响因子: 4
作者:
Schadt, EE;Li, C;Wong, WH
通讯作者: Wong, WH
DOI: 10.1186/gb-2003-4-6-r41
发表时间: 2003
期刊: Genome biology
影响因子: 12.3
作者:
Broberg P
通讯作者: Broberg P
DOI: 10.1093/bioinformatics/17.6.509
发表时间: 2001-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Baldi, P;Long, AD
通讯作者: Long, AD