How to decide? Different methods of calculating gene expression from short oligonucleotide array data will give different results.

How to decide? Different methods of calculating gene expression from short oligonucleotide array data will give different results.
复制标题

如何决定?从短寡核苷酸阵列数据中计算基因表达的不同方法将产生不同的结果。

DOI:
10.1186/1471-2105-7-137
复制
发表时间:
2006-03-15
期刊:
影响因子:
3
通讯作者:
Peeters, AJM
Peeters, AJM
中科院分区:
生物学4区
文献类型:
--
作者:
Millenaar, FF;Okyere, J;May, ST;van Zanten, M;Voesenek, LACJ;Peeters, AJM

文献摘要

参考文献

被引文献

相似文献

用于转录谱分析的短寡核苷酸阵列已经可用了几年。一般来说,来自这些阵列的原始数据借助于来自Affytechnology的微阵列分析套件或GeneChip操作软件(MAS或GCOS)进行分析。最近,有更多的方法来分析原始数据。理想情况下,所有这些方法应该或多或少得出相同的结果。我们开始评估不同的方法,并包括对我们自己的数据集的工作,以测试哪种方法给出最可靠的结果。使用相同(拟南芥)数据用6种不同算法(MAS 5、dChip PMMM、dChip PM、RMA、GC-RMA和PDNN)计算基因表达,导致不同的计算基因表达水平。因此,根据所使用的方法,不同的基因将被鉴定为差异调节。令人惊讶的是,不同方法之间只有27%到36%的重叠。此外,47.5%的基因/探针集显示出错配和完全匹配强度之间的良好相关性。在比较了六种算法后,RMA给出了最具重现性的结果,并显示出与所有方法鉴定为差异表达的基因的真实的时间RT-PCR数据的最高相关系数。然而,我们不能通过真实的时间RT-PCR验证大多数基因的微阵列结果,这些基因仅通过RMA计算。此外,我们得出的结论是,从完美匹配强度中减去失配强度很可能导致至少47.5%的表达值被显着低估。没有一个算法产生显著的表达值的基因存在的数量低于1 pmol。如果微阵列实验的唯一目的是发现新的候选基因,而发现的基因太多,那么通过对比方法预测的基因的互斥可以将新候选基因的列表缩小64%到73%。
Short oligonucleotide arrays for transcript profiling have been available for several years. Generally, raw data from these arrays are analysed with the aid of the Microarray Analysis Suite or GeneChip Operating Software (MAS or GCOS) from Affymetrix. Recently, more methods to analyse the raw data have become available. Ideally all these methods should come up with more or less the same results. We set out to evaluate the different methods and include work on our own data set, in order to test which method gives the most reliable results. Calculating gene expression with 6 different algorithms (MAS5, dChip PMMM, dChip PM, RMA, GC-RMA and PDNN) using the same (Arabidopsis) data, results in different calculated gene expression levels. Consequently, depending on the method used, different genes will be identified as differentially regulated. Surprisingly, there was only 27 to 36% overlap between the different methods. Furthermore, 47.5% of the genes/probe sets showed good correlation between the mismatch and perfect match intensities. After comparing six algorithms, RMA gave the most reproducible results and showed the highest correlation coefficients with Real Time RT-PCR data on genes identified as differentially expressed by all methods. However, we were not able to verify, by Real Time RT-PCR, the microarray results for most genes that were solely calculated by RMA. Furthermore, we conclude that subtraction of the mismatch intensity from the perfect match intensity results most likely in a significant underestimation for at least 47.5% of the expression values. Not one algorithm produced significant expression values for genes present in quantities below 1 pmol. If the only purpose of the microarray experiment is to find new candidate genes, and too many genes are found, then mutual exclusion of the genes predicted by contrasting methods can be used to narrow down the list of new candidate genes by 64 to 73%.
DOI: 10.1104/pp.104.051300
发表时间: 2005-02-01
期刊: PLANT PHYSIOLOGY
影响因子: 7.4
作者:
Allemeersch, J;Durinck, S;Kuiper, MTR
通讯作者: Kuiper, MTR
DOI: 10.1073/pnas.011404098
发表时间: 2001-01-02
影响因子: 11.1
作者:
Li, C;Wong, WH
通讯作者: Wong, WH
DOI: 10.1093/bioinformatics/btg410
发表时间: 2004-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cope, LM;Irizarry, RA;Speed, TP
通讯作者: Speed, TP
DOI: 10.1002/jcb.10073
发表时间: 2001-01-01
影响因子: 4
作者:
Schadt, EE;Li, C;Wong, WH
通讯作者: Wong, WH
DOI: 10.1104/pp.009415
发表时间: 2002-11-01
期刊: PLANT PHYSIOLOGY
影响因子: 7.4
作者:
Leclercq, J;Adams-Phillips, LC;Bouzayen, M
通讯作者: Bouzayen, M