Normalization Method for Transcriptional Studies of Heterogeneous Samples - Simultaneous Array Normalization and Identification of Equivalent Expression

Normalization Method for Transcriptional Studies of Heterogeneous Samples - Simultaneous Array Normalization and Identification of Equivalent Expression
复制标题

DOI:
10.2202/1544-6115.1339
复制
发表时间:
2009-01-01
影响因子:
0.9
通讯作者:
Satagopan, Jaya M.
Satagopan, Jaya M.
中科院分区:
数学4区
文献类型:
--
作者:
Qin, Li-Xuan;Satagopan, Jaya M.

文献摘要

被引文献

相似文献

归一化是转录图谱微阵列数据分析中的一个重要步骤,因为任何转录图谱实验中涉及的多个步骤往往会引起系统性的非生物变异。现有的数据归一化方法通常假设有很少或对称的差分表达式,但这一假设并不总是成立的。或者,非差异表达基因可用于阵列归一化。然而,从一开始就不知道哪些基因是非差异表达的。在本文中,我们提出了一个分层混合模型框架,以同时识别非差异表达基因并使用这些基因对阵列进行归一化。推导了对应于阵列效应的Fisher信息矩阵,为阵列归一化方法的选择提供了直观的指导。利用仿真数据对该方法的运行特性进行了评估。在各种参数配置下的仿真结果表明,该方法为阵列归一化提供了一种有用的选择。例如,在差异表达基因的适度流行以及过度表达和低表达程度不同的情况下,所提出的方法具有比中值归一化更好的敏感性。此外,当差异表达基因的流行率很小时,所提出的方法具有类似于中值归一化的性质。使用MSKCC的脂肪肉瘤研究来识别正常脂肪组织和脂肪肉瘤组织样本之间的差异表达基因,提供了所建议的方法的经验例证。
Normalization is an important step in the analysis of microarray data of transcription profiles as systematic non-biological variations often arise from the multiple steps involved in any transcription profiling experiment. Existing methods for data normalization often assume that there are few or symmetric differential expression, but this assumption does not always hold. Alternatively, non-differentially expressed genes may be used for array normalization. However, it is unknown at the outset which genes are non-differentially expressed. In this paper we propose a hierarchical mixture model framework to simultaneously identify non-differentially expressed genes and normalize arrays using these genes. The Fisher's information matrix corresponding to array effects is derived, which provides useful intuition for guiding the choice of array normalization method. The operating characteristics of the proposed method are evaluated using simulated data. The simulations conducted under a wide range of parametric configurations suggest that the proposed method provides a useful alternative for array normalization. For example, the proposed method has better sensitivity than median normalization under modest prevalence of differentially expressed genes and when the magnitudes of over-expression and under-expression are not the same. Further, the proposed method has properties similar to median normalization when the prevalence of differentially expressed genes is very small. Empirical illustration of the proposed method is provided using a liposarcoma study from MSKCC to identify genes differentially expressed between normal fat tissue versus liposarcoma tissue samples.