Batch effect removal methods for microarray gene expression data integration: a survey

Batch effect removal methods for microarray gene expression data integration: a survey
复制标题

DOI:
10.1093/bib/bbs037
复制
发表时间:
2013-07-01
影响因子:
9.5
通讯作者:
Nowe, Ann
Nowe, Ann
中科院分区:
生物学2区
文献类型:
--
作者:
Lazar, Cosmin;Meganck, Stijn;Nowe, Ann

文献摘要

被引文献

相似文献

基因组数据集成是实现大规模基因组数据分析的关键目标。由于基因组学实验产生的信息来源多种多样,这一过程非常具有挑战性。在这项工作中,我们回顾了设计来结合从微阵列基因表达(MAGE)实验记录的基因组数据的方法。人们已经认识到,不同MAGE数据集之间差异的主要来源是所谓的批处理效应。这里回顾的方法通过移除(或者更准确地说,尝试移除)与Batch Effect相关联的不需要的变化来执行数据集成。它们在一个统一的框架中提出,并附有广泛的评价工具,这些工具在评估数据整合过程的效率和质量时是强制性的。我们对MAGE数据集成方法进行了系统的描述,并提出了一些基本建议,以帮助用户选择合适的工具来集成MAGE数据进行大规模分析,以及如何从不同的角度对其进行评估,以量化其效率。本研究中用于说明目的的所有基因组数据均从InSilicoDB ext-LINK-TYPE=“URI”xlink:HREF=“http://insilico.ulb.ac.be”检索
Genomic data integration is a key goal to be achieved towards large-scale genomic data analysis. This process is very challenging due to the diverse sources of information resulting from genomics experiments. In this work, we review methods designed to combine genomic data recorded from microarray gene expression (MAGE) experiments. It has been acknowledged that the main source of variation between different MAGE datasets is due to the so-called 'batch effects'. The methods reviewed here perform data integration by removing (or more precisely attempting to remove) the unwanted variation associated with batch effects. They are presented in a unified framework together with a wide range of evaluation tools, which are mandatory in assessing the efficiency and the quality of the data integration process. We provide a systematic description of the MAGE data integration methodology together with some basic recommendation to help the users in choosing the appropriate tools to integrate MAGE data for large-scale analysis; and also how to evaluate them from different perspectives in order to quantify their efficiency. All genomic data used in this study for illustration purposes were retrieved from InSilicoDB ext-link-type="uri" xlink:href="http://insilico.ulb.ac.be"