Correlation test to assess low-level processing of high-density oligonucleotide microarray data.

Correlation test to assess low-level processing of high-density oligonucleotide microarray data.
复制标题

DOI:
10.1186/1471-2105-6-80
复制
发表时间:
2005-03-31
期刊:
影响因子:
3
通讯作者:
Pawitan Y
Pawitan Y
中科院分区:
生物学4区
文献类型:
--
作者:
Ploner A;Miller LD;Hall P;Bergh J;Pawitan Y

文献摘要

参考文献

被引文献

相似文献

目前,寡核苷酸阵列数据的低水平处理存在多种相互竞争的技术。技术的选择对后续的统计分析有深远影响,但在不参考外部数据的情况下,没有方法来评估某种特定技术是否适合某一特定数据集。 我们分析了基因之间的共调控,以检测阵列之间归一化不足的情况,其中共调控是通过统计相关性来衡量的。在大量基因中,随机选取的一对基因平均应具有零相关性,因此可以进行相关性检验。对于我们评估的所有数据集,以及包括MAS5、RMA和MBEI在内的三种最常用的低水平处理程序,管家基因归一化未通过检验。对于一个真实的临床数据集,RMA和MBEI显示缺失基因存在显著相关性。我们还发现,在探针组水平上进行的第二轮归一化在整体上显著改善了归一化效果。 文献中先前对低水平处理的评估仅限于人工掺入和混合数据集。在缺乏已知金标准的情况下,相关性标准使我们能够评估特定数据集低水平处理的适当性以及基因子集归一化的成功与否。
There are currently a number of competing techniques for low-level processing of oligonucleotide array data. The choice of technique has a profound effect on subsequent statistical analyses, but there is no method to assess whether a particular technique is appropriate for a specific data set, without reference to external data. We analyzed coregulation between genes in order to detect insufficient normalization between arrays, where coregulation is measured in terms of statistical correlation. In a large collection of genes, a random pair of genes should have on average zero correlation, hence allowing a correlation test. For all data sets that we evaluated, and the three most commonly used low-level processing procedures including MAS5, RMA and MBEI, the housekeeping-gene normalization failed the test. For a real clinical data set, RMA and MBEI showed significant correlation for absent genes. We also found that a second round of normalization on the probe set level improved normalization significantly throughout. Previous evaluation of low-level processing in the literature has been limited to artificial spike-in and mixture data sets. In the absence of a known gold-standard, the correlation criterion allows us to assess the appropriateness of low-level processing of a specific data set and the success of normalization for subsets of genes.
DOI: 10.1073/pnas.011404098
发表时间: 2001-01-02
影响因子: 11.1
作者:
Li, C;Wong, WH
通讯作者: Wong, WH
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1093/bioinformatics/btg410
发表时间: 2004-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cope, LM;Irizarry, RA;Speed, TP
通讯作者: Speed, TP
DOI: 10.1093/biostatistics/4.2.249
发表时间: 2003-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Irizarry, RA;Hobbs, B;Speed, TP
通讯作者: Speed, TP
DOI: 10.1186/gb-2005-6-2-r16
发表时间: 2005
期刊: Genome biology
影响因子: 12.3
作者:
Choe SE;Boutros M;Michelson AM;Church GM;Halfon MS
通讯作者: Halfon MS