A Comprehensive Evaluation of Generalizability of Deep Learning-Based Hi-C Resolution Improvement Methods.

A Comprehensive Evaluation of Generalizability of Deep Learning-Based Hi-C Resolution Improvement Methods.
复制标题

DOI:
10.3390/genes15010054
复制
发表时间:
2023-12-29
期刊:
影响因子:
3.5
通讯作者:
--
中科院分区:
生物学3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

Hi-C是一种广泛用于研究基因组3D结构的技术。由于其高测序成本,大多数生成的数据集是一个粗略的分辨率,这使得它不切实际的研究更精细的染色质特征,如拓扑相关结构域(TADs)和染色质环。最近提出了多种基于深度学习的方法,通过估算Hi-C读数(通常称为升级)来提高这些数据集的分辨率。然而,现有的工作评估这些方法的合成下采样数据集,或实验生成的稀疏Hi-C数据集的一个小子集,很难建立其在现实世界中的用例的推广。我们提出了我们的框架-Hi-CY-比较现有的Hi-C分辨率放大方法在七个实验生成的低分辨率Hi-C数据集属于不同级别的读取稀疏性源自三个细胞系的一套全面的评估指标。Hi-CY还包括四个下游分析任务,如重复序列和染色质环回忆,以提供这些方法的普遍性的全面报告。我们观察到现有的深度学习方法无法推广到实验生成的稀疏Hi-C数据集,性能下降高达57%。作为一种潜在的解决方案,我们发现使用实验生成的Hi-C数据集重新训练基于深度学习的方法可以将性能提高高达31%。更重要的是,Hi-CY表明,即使经过再训练,现有的基于深度学习的方法在提供稀疏Hi-C数据集时也很难恢复生物特征,如染色质环和TADs。通过Hi-CY框架,我们的研究强调了未来严格评估的必要性。我们确定了改进当前基于深度学习的Hi-C升级方法的具体途径,包括但不限于使用实验生成的数据集进行训练。
Hi-C is a widely used technique to study the 3D organization of the genome. Due to its high sequencing cost, most of the generated datasets are of a coarse resolution, which makes it impractical to study finer chromatin features such as Topologically Associating Domains (TADs) and chromatin loops. Multiple deep learning-based methods have recently been proposed to increase the resolution of these datasets by imputing Hi-C reads (typically called upscaling). However, the existing works evaluate these methods on either synthetically downsampled datasets, or a small subset of experimentally generated sparse Hi-C datasets, making it hard to establish their generalizability in the real-world use case. We present our framework—Hi-CY—that compares existing Hi-C resolution upscaling methods on seven experimentally generated low-resolution Hi-C datasets belonging to various levels of read sparsities originating from three cell lines on a comprehensive set of evaluation metrics. Hi-CY also includes four downstream analysis tasks, such as TAD and chromatin loops recall, to provide a thorough report on the generalizability of these methods. We observe that existing deep learning methods fail to generalize to experimentally generated sparse Hi-C datasets, showing a performance reduction of up to 57%. As a potential solution, we find that retraining deep learning-based methods with experimentally generated Hi-C datasets improves performance by up to 31%. More importantly, Hi-CY shows that even with retraining, the existing deep learning-based methods struggle to recover biological features such as chromatin loops and TADs when provided with sparse Hi-C datasets. Our study, through the Hi-CY framework, highlights the need for rigorous evaluation in the future. We identify specific avenues for improvements in the current deep learning-based Hi-C upscaling methods, including but not limited to using experimentally generated datasets for training.
DOI: 10.1038/nature12644
发表时间: 2013-11-14
期刊: NATURE
影响因子: 64.8
作者:
Jin, Fulai;Li, Yan;Dixon, Jesse R.;Selvaraj, Siddarth;Ye, Zhen;Lee, Ah Young;Yen, Chia-An;Schmitt, Anthony D.;Espinoza, Celso A.;Ren, Bing
通讯作者: Ren, Bing
DOI: 10.1186/s12864-018-4546-8
发表时间: 2018-02-23
期刊: BMC genomics
影响因子: 4.4
作者:
Oluwadare O;Zhang Y;Cheng J
通讯作者: Cheng J
DOI: 10.1038/s41467-018-03113-2
发表时间: 2018-02-21
影响因子: 16.6
作者:
Zhang Y;An L;Xu J;Zhang B;Zheng WJ;Hu M;Tang J;Yue F
通讯作者: Yue F
DeepHiC:用于增强 Hi-C 数据分辨率的生成对抗网络
DOI: 10.1371/journal.pcbi.1007287
发表时间: 2020-02-01
影响因子: 4.3
作者:
Hong, Hao;Jiang, Shuai;Bo, Xiaochen
通讯作者: Bo, Xiaochen
DOI: 10.1073/pnas.1911708117
发表时间: 2020-01-28
影响因子: 11.1
作者:
Pugacheva, Elena M.;Kubo, Naoki;Lobanenkov, Victor V.
通讯作者: Lobanenkov, Victor V.