Comparative Analysis of Transformation Methods for Gene Expression Profiles in Breast Cancer Datasets

Comparative Analysis of Transformation Methods for Gene Expression Profiles in Breast Cancer Datasets
复制标题

乳腺癌数据集中基因表达谱转化方法的比较分析

DOI:
--
复制
发表时间:
2016
期刊:
International Conferences on Biological Information and Biomedical Engineering
影响因子:
--
通讯作者:
H. Matsuda
H. Matsuda
中科院分区:
--
文献类型:
--
作者:
Yoshiaki Sota;S. Seno;Y. Takenaka;S. Noguchi;H. Matsuda

文献摘要

参考文献

相似文献

基因表达谱已越来越多地应用于临床实践。跨多个实验的表达数据的整合提供了对被检查的生物学的异质性的更好的洞察。数据集成的问题,来自平台或实验室来源的实验批次,仍然是系统地分析不同数据集数据的障碍。已经提出了几种方法(如ComBat)来消除批次效应。然而,这些方法通常假设基础数据的理想分布。当比较具有根本不同(依赖于数据集)分布的数据集时,可能会遇到困难。例如,临床数据集通常从具有各种疾病阶段或病症的患者样本收集。因此,我们在许多数据集上比较了几种数学变换,包括我们提出用于临床的非参数Z缩放变换方法(NPZ)。我们从Gene Expression Omnibus数据库中的24个Affytechnic HG-U133(GPL 96)或Affytechnic HG-U133 plus 2.0(GPL 570)数据集中选择了2,813例具有雌激素受体(ER)状态或人表皮生长因子受体2(HER 2)状态可用信息的患者。用以下四种方法之一处理微阵列表达数据:Raw(仅背景校正和对数转换)、Microarray Suite 5.0(MAS 5)、冷冻稳健多阵列分析(fRMA)和半径最小化(RMX)。通过使用以下五种方法之一对归一化数据进行顺序转换:未转换(无转换)、基于单阵列的转换(RANK、Z、NPZ或YuGene)。最后,我们比较了ER和HER 2的状态,通过免疫组织化学(IHC)染色与mRNA表达进行评估。我们发现,除了归一化之外,基于单阵列的转换提高了IHC染色的一致率。我们通过使用乳腺癌样本证明了转换的影响,并表明将基于单阵列的转换添加到微阵列表达数据中会导致与IHC染色的更强相关性。
Gene expression profiling has been increasingly used in clinical practice. Integration of expression data across multiple experiments provides better insight into the heterogeneity of the biology being examined. A problem of the data integration, an experimental batch from platform or laboratory sources, remains a barrier to systematically analyzing data across different datasets. Several methods (such as, ComBat) have been proposed to remove batch effects. However, these methods often make assumptions about ideal distribution of the underlying data. Difficulties might be expected when comparing datasets that have fundamentally different (dataset-dependent) distributions. For example, clinical datasets are often collected from patient samples with various disease stages or conditions. Therefore, we have compared several mathematical transformations across many datasets, including the nonparametric Z scaling transformation method (NPZ) we have proposed for clinical use. We selected 2,813 patients with available information on estrogen receptor (ER) status or human epidermal growth factor receptor 2 (HER2) status from 24 Affymetrix HG-U133 (GPL96) or Affymetrix HG-U133 plus 2.0 (GPL570) datasets in the Gene Expression Omnibus database. The microarray expression data were processed with one of the four following methods: Raw (background correction and log transformation only), Microarray Suite 5.0 (MAS5), frozen robust multiarray analysis (fRMA), and radius minimax (RMX). The normalized data were sequentially transformed by using one of the following five methods: untransformed (without transformation), single-array-based transformations (RANK, Z, NPZ, or YuGene). Finally, we compared the ER and HER2 statuses assessed by immunohistochemical (IHC) staining with mRNA expression. We found that single-array-based transformation in addition to normalization improved the concordance rates of the IHC staining. We demonstrated the influence of transformation by using breast cancer samples and showed that adding single-array-based transformations to microarray expression data resulted in stronger correlations with IHC staining.
DOI: 10.1200/jco.2008.18.1370
发表时间: 2009-03-10
影响因子: 45.3
作者:
Parker, Joel S.;Mullins, Michael;Bernard, Philip S.
通讯作者: Bernard, Philip S.
DOI: 10.1093/biostatistics/kxj037
发表时间: 2007-01-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Johnson, W. Evan;Li, Cheng;Rabinovic, Ariel
通讯作者: Rabinovic, Ariel
DOI: 10.1093/biostatistics/kxp059
发表时间: 2010-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
McCall, Matthew N.;Bolstad, Benjamin M.;Irizarry, Rafael A.
通讯作者: Irizarry, Rafael A.