Comparison of normalization methods for Hi-C data

Comparison of normalization methods for Hi-C data
复制标题

Hi-C 数据标准化方法比较

DOI:
10.2144/btn-2019-0105
复制
发表时间:
2020-02-01
期刊:
影响因子:
2.7
通讯作者:
Wu, Zhifang
Wu, Zhifang
中科院分区:
工程技术4区
文献类型:
--
作者:
Lyu, Hongqiang;Liu, Erhu;Wu, Zhifang

文献摘要

被引文献

相似文献

HI-C主要用于研究基因组间的相互作用。在Hi-C实验中,人们认为源于不同系统偏差的偏差会导致原始样本之间的外部变异性,并影响下游解释的可靠性。归一化作为Hi-C分析中的一条重要途径,旨在消除不需要的系统性偏差;因此,比较Hi-C归一化方法有利于它们的选择和下游分析。本文从多方面考虑,对6种Hi-C归一化方法进行了综合比较。根据比较结果,已经表明,在大多数考虑因素中,交叉样本方法明显优于单独样本方法。分析了这几种归一化方法的差异,给出了一些实用的建议,并以表格的形式对结果进行了总结,以便于选择六种归一化方法。从热图纹理、统计质量、分辨率的影响、距离层的一致性和拓扑关联区域结构的可重复性等多个方面对6种Hi-C数据归一化方法进行了综合比较,这些方法的实现源代码可在https://github.com/lhqxinghun/bioinformatics/tree/master/Hi-C/NormCompareMETHOD摘要中找到。在这些考虑中,从交互频率的分布、副本的相关性和上下文之间副本的可比性三个方面对统计质量进行了深入的研究。并对这些方法的性能进行了比较。
Hi-C has been predominately used to study the genome-wide interactions of genomes. In Hi-C experiments, it is believed that biases originating from different systematic deviations lead to extraneous variability among raw samples, and affect the reliability of downstream interpretations. As an important pipeline in Hi-C analysis, normalization seeks to remove the unwanted systematic biases; thus, a comparison between Hi-C normalization methods benefits their choice and the downstream analysis. In this article, a comprehensive comparison is proposed to investigate six Hi-C normalization methods in terms of multiple considerations. In light of comparison results, it has been shown that a cross-sample approach significantly outperforms individual sample methods in most considerations. The differences between these methods are analyzed, some practical recommendations are given, and the results are summarized in a table to facilitate the choice of the six normalization methods. The source code for the implementation of these methods is available at https://github.com/lhqxinghun/bioinformatics/tree/master/Hi-C/NormCompareMETHOD SUMMARY Six normalization methods for Hi-C data were compared comprehensively in terms of multiple considerations, including heat map texture, statistical quality, influence of resolution, consistency of distance stratum and reproducibility of topologically associating domain architecture. Among these considerations, the quality of statistics was investigated in depth from three aspects, comprising distribution of interaction frequency, correlation of replicates and comparability of replicates between contexts. The performance of these methods is compared.