Inference of 3D genome architecture by modeling overdispersion of Hi-C data.

Inference of 3D genome architecture by modeling overdispersion of Hi-C data.
复制标题

DOI:
10.1093/bioinformatics/btac838
复制
发表时间:
2023-01-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

我们解决了从Hi-C数据推断基因组结构的共识3D模型的挑战。现有的方法通常依赖于两步算法:首先,将接触计数转换为距离,然后优化类似于多维缩放(MDS)的目标函数以推断3D模型。其他方法使用最大似然方法,将两个位点之间的接触计数建模为泊松随机变量,其强度是它们之间的距离的递减函数。然而,接触计数的泊松模型意味着数据的方差等于平均值,这种关系通常过于限制,无法正确建模计数数据。我们首先确认存在的过度分散在几个真实的Hi-C数据集,我们表明,过度分散甚至出现在模拟数据集。然后,我们提出了一个新的模型,称为Pastis-NB,在那里我们取代泊松模型的接触计数的负二项,这是参数化的平均值和一个单独的分散参数。离差参数允许独立于均值调整方差,从而更好地对过度分散的数据进行建模。我们比较的结果Pastis-NB的几个以前发表的算法,MDS为基础的和统计方法。我们表明,负二项推理在模拟数据上产生了更准确的结构,并且在真实的Hi-C重复和不同分辨率下比其他模型产生了更稳健的结构。Pastis-NB的Python实现可以在BSD许可证下在https://github.com/hiclib/pastis上获得。 补充数据可在Bioinformatics在线获得。
We address the challenge of inferring a consensus 3D model of genome architecture from Hi-C data. Existing approaches most often rely on a two-step algorithm: first, convert the contact counts into distances, then optimize an objective function akin to multidimensional scaling (MDS) to infer a 3D model. Other approaches use a maximum likelihood approach, modeling the contact counts between two loci as a Poisson random variable whose intensity is a decreasing function of the distance between them. However, a Poisson model of contact counts implies that the variance of the data is equal to the mean, a relationship that is often too restrictive to properly model count data. We first confirm the presence of overdispersion in several real Hi-C datasets, and we show that the overdispersion arises even in simulated datasets. We then propose a new model, called Pastis-NB, where we replace the Poisson model of contact counts by a negative binomial one, which is parametrized by a mean and a separate dispersion parameter. The dispersion parameter allows the variance to be adjusted independently from the mean, thus better modeling overdispersed data. We compare the results of Pastis-NB to those of several previously published algorithms, both MDS-based and statistical methods. We show that the negative binomial inference yields more accurate structures on simulated data, and more robust structures than other models across real Hi-C replicates and across different resolutions. A Python implementation of Pastis-NB is available at https://github.com/hiclib/pastis under the BSD license. Supplementary data are available at Bioinformatics online.
DOI: 10.1038/nature12644
发表时间: 2013-11-14
期刊: NATURE
影响因子: 64.8
作者:
Jin, Fulai;Li, Yan;Dixon, Jesse R.;Selvaraj, Siddarth;Ye, Zhen;Lee, Ah Young;Yen, Chia-An;Schmitt, Anthony D.;Espinoza, Celso A.;Ren, Bing
通讯作者: Ren, Bing
DOI: 10.1038/nbt.2057
发表时间: 2011-12-25
影响因子: 46.9
作者:
通讯作者: --
用于分析HI-C数据的二维分割。
DOI: 10.1093/bioinformatics/btu443
发表时间: 2014-09-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lévy-Leduc C;Delattre M;Mary-Huard T;Robin S
通讯作者: Robin S
DOI: 10.1101/gr.169417.113
发表时间: 2014-06
期刊: Genome research
影响因子: 7
作者:
Ay F;Bunnik EM;Varoquaux N;Bol SM;Prudhomme J;Vert JP;Noble WS;Le Roch KG
通讯作者: Le Roch KG
DOI: 10.1007/bf01589116
发表时间: 1989-12-01
影响因子: 2.7
作者:
LIU, DC;NOCEDAL, J
通讯作者: NOCEDAL, J