Unsupervised Contrastive Peak Caller for ATAC-seq.

Unsupervised Contrastive Peak Caller for ATAC-seq.
复制标题

用于 ATAC-seq 的无监督对比峰值调用器。

DOI:
10.1101/2023.01.07.523108
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Dorman,Karin
Dorman,Karin
中科院分区:
--
文献类型:
--
作者:
Vu,HaTH;Zhang,Yudi;Tuteja,Geetu;Dorman,Karin

文献摘要

相似文献

转座酶可及染色质测序试验(ATAC-seq)是一种常见的试验,通过使用Tn 5转座酶鉴定染色质可及区域,该转座酶可接近、切割衔接子并将衔接子连接至DNA片段,用于随后的扩增和测序。在称为“峰识别”的过程中对这些测序区域进行定量和富集测试。大多数无监督的峰调用方法基于简单的统计模型,并且具有较高的假阳性率。新开发的监督式深度学习方法可能会取得成功,但它们依赖于高质量的标记数据进行训练,而这些数据可能很难获得。此外,尽管生物复制被认为是重要的,但在深度学习工具中使用复制没有既定的方法,并且传统方法可用的方法要么不能应用于ATAC-seq,其中对照样品可能不可用,要么是事后的,并且不利用读段富集数据中潜在复杂但可再现的信号。在这里,我们提出了一种新的峰值调用器,它使用无监督的对比学习来从多个重复中提取共享信号。对原始覆盖数据进行编码以获得低维嵌入,并进行优化以最小化生物重复的对比损失。这些嵌入被传递到另一个对比损失,用于学习和预测峰值,并在自动编码器损失下解码为去噪数据。我们将我们的复制对比学习器(RCL)方法与ATAC-seq数据上的其他现有方法进行了比较,使用来自ChromHMM基因组标签和转录因子ChIP-seq的注释作为噪声真理。RCL始终取得最佳业绩。
The assay for transposase-accessible chromatin with sequencing (ATAC-seq) is a common assay to identify chromatin accessible regions by using a Tn5 transposase that can access, cut, and ligate adapters to DNA fragments for subsequent amplification and sequencing. These sequenced regions are quantified and tested for enrichment in a process referred to as “peak calling.” Most unsupervised peak calling methods are based on simple statistical models and suffer from elevated false positive rates. Newly developed supervised deep learning methods can be successful, but they rely on high quality labeled data for training, which can be difficult to obtain. Moreover, though biological replicates are recognized to be important, there are no established approaches for using replicates in the deep learning tools, and the approaches available for traditional methods either cannot be applied to ATAC-seq, where control samples may be unavailable, or are post hoc and do not capitalize on potentially complex, but reproducible signal in the read enrichment data. Here, we propose a novel peak caller that uses unsupervised contrastive learning to extract shared signals from multiple replicates. Raw coverage data are encoded to obtain low-dimensional embeddings and optimized to minimize a contrastive loss over biological replicates. These embeddings are passed to another contrastive loss for learning and predicting peaks and decoded to denoised data under an autoencoder loss. We compared our replicative contrastive learner (RCL) method with other existing methods on ATAC-seq data, using annotations from ChromHMM genomic labels and transcription factor ChIP-seq as noisy truth. RCL consistently achieved the best performance.