Assessing the need for sequence-based normalization in tiling microarray experiments

Assessing the need for sequence-based normalization in tiling microarray experiments
复制标题

DOI:
10.1093/bioinformatics/btm052
复制
发表时间:
2007-04-15
期刊:
影响因子:
5.8
通讯作者:
Gerstein, Mark B.
Gerstein, Mark B.
中科院分区:
生物学3区
文献类型:
--
作者:
Royce, Thomas E.;Rozowsky, Joel S.;Gerstein, Mark B.

文献摘要

被引文献

相似文献

动机:微阵列特征密度的增加使得所谓的拼接微阵列得以构建。这些阵列或阵列组包含以规则的基因组间隔靶向已测序基因组区域的探针。这种方法的无偏性特性使得能够鉴定新的转录序列、定位转录因子结合位点(染色质免疫沉淀 - 芯片)以及进行高分辨率比较基因组杂交等应用。随着拼接微阵列变得更加经济实惠,这些应用迅速普及。为了实现最大效用,拼接微阵列平台需要发展到能够实现1个核苷酸分辨率,并且我们对在如此精细分辨率下进行的单个测量有信心。必须系统地消除拼接阵列信号中的任何偏差以实现这一目标。 结果:为此,我们研究了探针序列组成对拼接微阵列用于鉴定新的转录和转录因子结合位点的功效的重要性。我们发现强度高度依赖于序列,并且会极大地影响结果。我们开发了三种用于评估这种序列依赖性的指标,并将它们用于评估拼接微阵列文献中现有的基于序列的归一化方法。此外,我们应用了三种新技术来解决这个问题;一种方法是从基因芯片品牌微阵列的类似工作改编而来,它基于将阵列信号建模为探针序列的线性函数,第二种方法通过对模型进行迭代加权和重新拟合扩展了这种方法,第三种技术将用于阵列间归一化的流行的分位数归一化算法外推到探针序列空间。根据此处定义的指标,这三种方法相较于现有策略表现良好。
Motivation: Increases in microarray feature density allow the construction of so-called tiling microarrays. These arrays, or sets of arrays, contain probes targeting regions of sequenced genomes at regular genomic intervals. The unbiased nature of this approach allows for the identification of novel transcribed sequences, the localization of transcription factor binding sites (ChIP-chip), and high resolution comparative genomic hybridization, among other uses. These applications are quickly growing in popularity as tiling microarrays become more affordable. To reach maximum utility, the tiling microarray platform needs be developed to the point that 1 nt resolutions are achieved and that we have confidence in individual measurements taken at this fine of resolution. Any biases in tiling array signals must be systematically removed to achieve this goal.Results: Towards this end, we investigated the importance of probe sequence composition on the efficacy of tiling microarrays for identifying novel transcription and transcription factor binding sites. We found that intensities are highly sequence dependent and can greatly influence results. We developed three metrics for assessing this sequence dependence and use them in evaluating existing sequence-based normalizations from the tiling microarray literature. In addition, we applied three new techniques for addressing this problem; one method, adapted from similar work on GeneChip brand microarrays, is based on modeling array signal as a linear function of probe sequence, the second method extends this approach by iterative weighting and re-fitting of the model, and the third technique extrapolates the popular quantile normalization algorithm for between-array normalization to probe sequence space. These three methods perform favorably to existing strategies, based on the metrics defined here.