Normalization of high dimensional genomics data where the distribution of the altered variables is skewed.

Normalization of high dimensional genomics data where the distribution of the altered variables is skewed.
复制标题

DOI:
10.1371/journal.pone.0027942
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Stenberg P
Stenberg P
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Landfors M;Philip P;Rydén P;Stenberg P

文献摘要

参考文献

被引文献

相似文献

现在常规地使用基于不同阵列或测序的技术进行基因表达或蛋白质结合模式的全基因组分析,以比较不同人群,例如治疗组和参考组。通常有必要对获得的数据进行标准化,以消除在进行实验工作的过程中引入的技术变化,但是标准的标准化技术不能消除真正改变的变量的分布是偏斜的情况下的技术偏差,即当大部分变量受到治疗的积极或消极影响时。然而,几个实验可能会产生这样的偏态分布,包括用于染色质研究的ChIP芯片实验,用于细胞凋亡研究的基因表达实验,以及正常和肿瘤组织中拷贝数变异的SNP研究。一个初步的研究,使用穗在阵列数据建立的能力,一个实验,以确定改变的变量,并产生无偏估计的倍数变化减少的改变变量的分数和偏度的增加。我们提出了以下的工作流程来分析高维实验与区域的变量改变:(1)预处理原始数据使用标准的归一化技术之一。(2)调查改变的变量的分布是否偏斜。(3)如果不认为分布是偏斜的,则不需要额外的归一化。否则,使用新的HMM辅助归一化过程重新归一化数据。(4)进行下游分析。在这里,ChIP芯片数据和模拟数据被用来评估工作流的性能。研究发现,偏态分布可以通过使用新的DSE检验(检测偏态实验)检测。此外,将HMM辅助的归一化应用于其中真正改变的变量的分布是偏斜的实验,结果比使用标准和不变归一化方法可以获得的灵敏度高得多,偏差低得多。
Genome-wide analysis of gene expression or protein binding patterns using different array or sequencing based technologies is now routinely performed to compare different populations, such as treatment and reference groups. It is often necessary to normalize the data obtained to remove technical variation introduced in the course of conducting experimental work, but standard normalization techniques are not capable of eliminating technical bias in cases where the distribution of the truly altered variables is skewed, i.e. when a large fraction of the variables are either positively or negatively affected by the treatment. However, several experiments are likely to generate such skewed distributions, including ChIP-chip experiments for the study of chromatin, gene expression experiments for the study of apoptosis, and SNP-studies of copy number variation in normal and tumour tissues. A preliminary study using spike-in array data established that the capacity of an experiment to identify altered variables and generate unbiased estimates of the fold change decreases as the fraction of altered variables and the skewness increases. We propose the following work-flow for analyzing high-dimensional experiments with regions of altered variables: (1) Pre-process raw data using one of the standard normalization techniques. (2) Investigate if the distribution of the altered variables is skewed. (3) If the distribution is not believed to be skewed, no additional normalization is needed. Otherwise, re-normalize the data using a novel HMM-assisted normalization procedure. (4) Perform downstream analysis. Here, ChIP-chip data and simulated data were used to evaluate the performance of the work-flow. It was found that skewed distributions can be detected by using the novel DSE-test (Detection of Skewed Experiments). Furthermore, applying the HMM-assisted normalization to experiments where the distribution of the truly altered variables is skewed results in considerably higher sensitivity and lower bias than can be attained using standard and invariant normalization methods.
DOI: 10.1371/journal.pgen.1000814
发表时间: 2010-01-15
期刊: PLoS genetics
影响因子: 4.5
作者:
Nègre N;Brown CD;Shah PK;Kheradpour P;Morrison CA;Henikoff JG;Feng X;Ahmad K;Russell S;White RA;Stein L;Henikoff S;Kellis M;White KP
通讯作者: White KP
DOI: 10.1186/1471-2105-11-503
发表时间: 2010-10-11
期刊: BMC bioinformatics
影响因子: 3
作者:
Freyhult E;Landfors M;Önskog J;Hvidsten TR;Rydén P
通讯作者: Rydén P
DOI: 10.1093/bioinformatics/17.6.509
发表时间: 2001-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Baldi, P;Long, AD
通讯作者: Long, AD
一种隐藏的马尔可夫模型方法,用于确定基因组瓷砖微阵列的表达。
DOI: 10.1186/1471-2105-7-239
发表时间: 2006-05-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Munch K;Gardner PP;Arctander P;Krogh A
通讯作者: Krogh A
DOI: 10.1371/journal.pbio.1000320
发表时间: 2010-02-23
期刊: PLoS biology
影响因子: 9.8
作者:
Zhang Y;Malone JH;Powell SK;Periwal V;Spana E;Macalpine DM;Oliver B
通讯作者: Oliver B