Quantification of histone modification ChIP-seq enrichment for data mining and machine learning applications.

Quantification of histone modification ChIP-seq enrichment for data mining and machine learning applications.
复制标题

DOI:
10.1186/1756-0500-4-288
复制
发表时间:
2011-08-11
期刊:
影响因子:
1.8
通讯作者:
Bekiranov S
Bekiranov S
中科院分区:
其他
文献类型:
--
作者:
Hoang SA;Xu X;Bekiranov S

文献摘要

被引文献

相似文献

ChIP-seq 技术的出现使表观遗传调控网络的研究成为一个计算上易于处理的问题。几个小组已将统计计算方法应用于 ChIP-seq 数据集,以深入了解转录的表观遗传调控。然而,用于估计这些计算研究的 ChIP-seq 数据富集水平的方法尚未得到充分研究且存在差异。由于从这些数据挖掘和机器学习应用程序中得出的结论在很大程度上取决于丰富水平输入,因此应该对统计模型性能的估计方法进行比较。使用各种方法来估计 20 个组蛋白甲基化和组蛋白变体 H2A.Z 的基因 ChIP-seq 富集水平。多元自适应回归样条(MARS)算法应用于每种估计方法,使用富集水平的估计作为预测因子,基因表达水平作为响应。用于估计富集水平的方法包括标签计数和应用于整个基因和特定基因区域的基于模型的方法。这些方法也适用于各种大小的估计窗口。 MARS 模型的性能通过广义交叉验证评分 (GCV) 进行评估。我们确定基于模型的富集估计方法(基于平均模式的空间权重富集)提供了对标签计数方法的改进。此外,包含整个基因体信息的方法比关注基因特定子区域(例如 5' 或 3' 区域)的方法有所改进。通过使用整个基因体的数据并结合富集的空间分布,可以提高数据挖掘和机器学习方法应用于组蛋白修饰 ChIP-seq 数据时的性能。富集度估计的细化最终提高了模型预测的准确性。
The advent of ChIP-seq technology has made the investigation of epigenetic regulatory networks a computationally tractable problem. Several groups have applied statistical computing methods to ChIP-seq datasets to gain insight into the epigenetic regulation of transcription. However, methods for estimating enrichment levels in ChIP-seq data for these computational studies are understudied and variable. Since the conclusions drawn from these data mining and machine learning applications strongly depend on the enrichment level inputs, a comparison of estimation methods with respect to the performance of statistical models should be made. Various methods were used to estimate the gene-wise ChIP-seq enrichment levels for 20 histone methylations and the histone variant H2A.Z. The Multivariate Adaptive Regression Splines (MARS) algorithm was applied for each estimation method using the estimation of enrichment levels as predictors and gene expression levels as responses. The methods used to estimate enrichment levels included tag counting and model-based methods that were applied to whole genes and specific gene regions. These methods were also applied to various sizes of estimation windows. The MARS model performance was assessed with the Generalized Cross-Validation Score (GCV). We determined that model-based methods of enrichment estimation that spatially weight enrichment based on average patterns provided an improvement over tag counting methods. Also, methods that included information across the entire gene body provided improvement over methods that focus on a specific sub-region of the gene (e.g., the 5' or 3' region). The performance of data mining and machine learning methods when applied to histone modification ChIP-seq data can be improved by using data across the entire gene body, and incorporating the spatial distribution of enrichment. Refinement of enrichment estimation ultimately improved accuracy of model predictions.