Combining gene expression data from different generations of oligonucleotide arrays.

Combining gene expression data from different generations of oligonucleotide arrays.
复制标题

DOI:
10.1186/1471-2105-5-159
复制
发表时间:
2004-10-25
期刊:
影响因子:
3
通讯作者:
Park PJ
Park PJ
中科院分区:
生物学4区
文献类型:
--
作者:
Hwang KB;Kong SW;Greenberg SA;Park PJ

文献摘要

参考文献

被引文献

相似文献

微阵列分析中的一个重要挑战是充分利用先前积累的数据,无论是来自自己的实验室还是来自公共存储库。通过对各种数据集的比较分析,可以更全面地了解潜在的机制或结构。然而,正如我们在这项工作中发现的那样,基因组序列注释和探针设计标准的不断变化使得即使来自相同微阵列平台的不同代的基因表达数据也难以进行比较。我们首先描述了聚类分析和差异表达基因鉴定中揭示的两代Affymetrix寡核苷酸阵列结果之间的不一致程度。然后,我们提出了一种方法来增加可比性。我们使用的数据集由一组来自炎性肌病患者的14个人类肌肉活检样本组成,这些样本在HG-U95 Av 2和HG-U133 A人类阵列上杂交。我们发现,使用探针组匹配表进行比较分析提供的Affyphase产生更好的结果比匹配UniGene或LocusLink标识符,但仍然不够。对样本中每个基因的表达值进行重新缩放以及通过表达值进行数据过滤增强了可比性,但仅适用于少数特定分析。作为提高可比性的通用方法,我们选择了两种阵列类型中具有重叠序列片段的探针子集,并仅基于所选探针重新计算表达值。我们表明,这种过滤的探针显着提高了可比性,同时保留了足够数量的探针集,以供进一步分析。高密度寡核苷酸阵列之间的兼容性受到探针水平序列信息的显著影响。通过基于它们的序列重叠仔细过滤探针,可以更有效地组合来自不同代微阵列的数据。
One of the important challenges in microarray analysis is to take full advantage of previously accumulated data, both from one's own laboratory and from public repositories. Through a comparative analysis on a variety of datasets, a more comprehensive view of the underlying mechanism or structure can be obtained. However, as we discover in this work, continual changes in genomic sequence annotations and probe design criteria make it difficult to compare gene expression data even from different generations of the same microarray platform. We first describe the extent of discordance between the results derived from two generations of Affymetrix oligonucleotide arrays, as revealed in cluster analysis and in identification of differentially expressed genes. We then propose a method for increasing comparability. The dataset we use consists of a set of 14 human muscle biopsy samples from patients with inflammatory myopathies that were hybridized on both HG-U95Av2 and HG-U133A human arrays. We find that the use of the probe set matching table for comparative analysis provided by Affymetrix produces better results than matching by UniGene or LocusLink identifiers but still remains inadequate. Rescaling of expression values for each gene across samples and data filtering by expression values enhance comparability but only for few specific analyses. As a generic method for improving comparability, we select a subset of probes with overlapping sequence segments in the two array types and recalculate expression values based only on the selected probes. We show that this filtering of probes significantly improves the comparability while retaining a sufficient number of probe sets for further analysis. Compatibility between high-density oligonucleotide arrays is significantly affected by probe-level sequence information. With a careful filtering of the probes based on their sequence overlaps, data from different generations of microarrays can be combined more effectively.
DOI: 10.1093/nar/29.5.e29
发表时间: 2001-03-01
影响因子: 14.9
作者:
Baugh, L. R.;Hill, A. A.;Hunter, Craig P.
通讯作者: Hunter, Craig P.
DOI: 10.1023/b:plan.0000019069.23317.97
发表时间: 2003-11-01
影响因子: 5.1
作者:
Hennig, L;Menges, M;Gruissem, W
通讯作者: Gruissem, W
DOI: 10.1073/pnas.1534744100
发表时间: 2003-09-30
影响因子: 11.1
作者:
Mei, R;Hubbell, E;Webster, TA
通讯作者: Webster, TA
DOI: 10.1186/1471-2164-4-31
发表时间: 2003-07-29
期刊: BMC genomics
影响因子: 4.4
作者:
Huminiecki L;Lloyd AT;Wolfe KH
通讯作者: Wolfe KH
DOI: 10.1101/gr.1048803
发表时间: 2003-07-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Barczak, A;Rodriguez, MW;Erle, DJ
通讯作者: Erle, DJ