Model-based analysis of oligonucleotide arrays: Expression index computation and outlier detection

Model-based analysis of oligonucleotide arrays: Expression index computation and outlier detection
复制标题

DOI:
10.1073/pnas.011404098
复制
发表时间:
2001-01-02
影响因子:
11.1
通讯作者:
Wong, WH
Wong, WH
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Li, C;Wong, WH

文献摘要

被引文献

相似文献

cDNA和寡核苷酸DNA阵列的最新进展使得同时测量许多基因的mRNA转录本的丰度成为可能。这种实验的分析是不平凡的,因为大的数据大小和许多层次的变化在不同阶段的实验。由于用于询问同一基因的不同探针之间可能存在巨大差异,该分析变得更加复杂。然而,一个有吸引力的特点,高密度寡核苷酸阵列,如那些通过光刻和喷墨技术生产的芯片制造和杂交过程的标准化。因此,探针特异性偏倚虽然显著,但具有高度可重复性和可预测性,并且可以通过适当的建模和分析方法来减少其不利影响。在这里,我们提出了一个统计模型的探针水平的数据,并开发基于模型的基因表达指数的估计。我们还提出了基于模型的方法,用于识别和处理交叉杂交探针和污染阵列区域。这些结果的应用将在其他地方介绍。
Recent advances in cDNA and oligonucleotide DNA arrays have made it possible to measure the abundance of mRNA transcripts for many genes simultaneously. The analysis of such experiments is nontrivial because of large data size and many levels of variation introduced at different stages of the experiments. The analysis is further complicated by the large differences that may exist among different probes used to interrogate the same gene. However, an attractive feature of high-density oligonucleotide arrays such as those produced by photolithography and inkjet technology is the standardization of chip manufacturing and hybridization process. As a result, probe-specific biases, although significant, are highly reproducible and predictable, and their adverse effect can be reduced by proper modeling and analysis methods. Here, we propose a statistical model for the probe-level data, and develop model-based estimates for gene expression indexes. We also present model-based methods for identifying and handling cross-hybridizing probes and contaminating array regions. Applications of these results will be presented elsewhere.