Large-scale imputation of epigenomic datasets for systematic annotation of diverse human tissues.

Large-scale imputation of epigenomic datasets for systematic annotation of diverse human tissues.
复制标题

DOI:
10.1038/nbt.3157
复制
发表时间:
2015-04
影响因子:
46.9
通讯作者:
Kellis M
Kellis M
中科院分区:
工程技术1区
文献类型:
--
作者:
Ernst J;Kellis M

文献摘要

被引文献

相似文献

有了数百个表观基因组图谱,就有机会利用标记和样本中表观遗传信号的相关性,对其他数据集进行大规模预测。在这里,我们通过回归树的集合利用这种相关性进行表观基因组imputation。我们计算了4315张高分辨率信号图,其中26%是实验观测到的。输入的信号轨迹与观察到的信号总体上相似,并且在一致性、基因注释的恢复和疾病相关变异的富集方面优于实验数据集。我们使用输入的数据来检测低质量的实验数据集,寻找具有意想不到的表观基因组信号的基因组位点,为新实验定义高优先级标记,并描述跨越不同组织和细胞类型的127个参考表观基因组的染色质状态。我们的输入数据集提供了迄今为止最全面的人类调控注释,我们的方法和ChromImpute软件构成了对表观基因组信息大规模实验制图的有用补充。
With hundreds of epigenomic maps, the opportunity arises to exploit the correlated nature of epigenetic signals, across both marks and samples, for large-scale prediction of additional datasets. Here, we undertake epigenome imputation by leveraging such correlations through an ensemble of regression trees. We impute 4,315 high-resolution signal maps, of which 26% are also experimentally observed. Imputed signal tracks show overall similarity to observed signals, and surpass experimental datasets in consistency, recovery of gene annotations, and enrichment for disease-associated variants. We use the imputed data to detect low quality experimental datasets, to find genomic sites with unexpected epigenomic signals, to define high-priority marks for new experiments, and to delineate chromatin states in 127 reference epigenomes spanning diverse tissues and cell types. Our imputed datasets provide the most comprehensive human regulatory annotation to date, and our approach and the ChromImpute software constitute a useful complement to large-scale experimental mapping of epigenomic information.