Filtering STARR-Seq Peaks for Enhancers with Sequence Models

Filtering STARR-Seq Peaks for Enhancers with Sequence Models
复制标题

使用序列模型过滤 STARR-Seq 峰以获取增强子

DOI:
10.1145/3388440.3414905
复制
发表时间:
2020
期刊:
Computational Biology and Health Informatics
影响因子:
--
通讯作者:
Halligan, Benjamin
Halligan, Benjamin
中科院分区:
--
文献类型:
--
作者:
Nowling, Ronald J.;Geromel, Rafael Reple;Halligan, Benjamin

文献摘要

相似文献

STARR-Seq是一种用于直接鉴定具有增强子活性的基因组区域的高通量技术[1]。剪切基因组DNA,将其插入设计成使得具有增强子活性的DNA触发自转录的人工质粒中,并转染到培养细胞中。将得到的RNA转化回cDNA,测序,并与参考基因组比对。通过使用统计检验将每个点处观察到的读取深度与对照DNA的预期读取深度进行比较来调用“峰”。基于读取深度的峰识别方法的实例包括MACS 2 [4]、basicSTARRSeq和STARRPeaker [3]。在平均读取深度低但方差高的区域中准确区分真实的峰和伪影具有挑战性。幸运的是,增强子活性与序列含量密切相关。我们建议在半监督框架中使用基于序列的机器学习模型来过滤峰值。从黑腹果蝇dm 3基因组中提取了以[1]中的111 k STARR峰为中心的501 bp序列。随机取样的501-bp序列用作阴性组。使用Bonferroni校正的显著性值(α = 0.05)对峰进行过滤,以创建2.2k个峰的“高置信度”子集。具有k-mer计数特征的逻辑回归模型在高置信度峰序列及其阴性上进行训练,并用于对剩余的约8.8k个峰序列进行分类。自我训练的基于序列的模型鉴定了另外的103.7k个候选增强子(“中等置信度”)。我们绘制了三组峰(高、中和低置信度)的读数深度对数倍数变化的直方图(见图1)。中置信度和低置信度峰的分布显著重叠。基于序列的模型识别了增强子候选者,否则这些候选者将单独使用读取深度被过滤掉。黑腹鼠FAIRE-Seq数据集来自[2]。用Trimmomatic清洗测序数据,用bwa回溯与dm 3基因组比对,并用samtools过滤作图质量(q < 10)。MACS 2称为61 k FAIRE峰值。STARR峰与FAIRE峰重叠,精密度分别为52.7%(高置信度峰)、40.6%(中等置信度峰)和22.5%(低置信度峰)。
STARR-Seq is a high-throughput technique for directly identifying genomic regions with enhancer activity [1]. Genomic DNA is sheared, inserted into artificial plasmids designed so that DNA with enhancer activity trigger self-transcription, and transfected into culture cells. The resulting RNA is converted back into cDNA, sequenced, and aligned to a reference genome. "Peaks" are called by comparing observed read depth at each point to an expected read depth from control DNA using a statistical test. Examples of peak calling methods based on read depth include MACS2 [4], basicSTARRSeq, and STARRPeaker [3].It is challenging to accurately distinguish between real peaks and artifacts in regions where mean read depth is low but the variance is high. Fortunately, enhancer activity is strongly correlated with sequence content. We propose using sequence-based machine learning models in a semi-supervised framework to filter peaks. 501-bp sequences centered on the ≈11k STARR peaks from [1] were extracted from the Drosophila melanogaster dm3 genome. Randomly-sampled 501-bp sequences were used as a negative set. Peaks were filtered using a Bonferroni-corrected significance value (α = 0.05) to create a "high-confidence" subset of ≈2.2k peaks. A Logistic Regression model with k-mer count features was trained on the high-confidence peak sequences and their negatives and used to classifying the remaining ≈8.8k peak sequences. The self-trained, sequenced-based model identified an additional ≈3.7k candidate enhancers ("medium confidence"). The remaining ≈5k STARR peaks were considered "low confidence" peaks.We plotted histograms of the read depth log-fold change for the three sets of peaks (high, medium, and low confidence) (see Figure 1). The distributions for the medium- and low-confidence peaks overlapped significantly. The sequence-based model identified enhancer candidates that would otherwise be filtered out using read depth alone.We called peaks for the 4 D. melanogaster FAIRE-Seq data sets from [2]. Sequencing data were cleaned with Trimmomatic, aligned to the dm3 genome with bwa backtrack, and filtered for mapping quality (q < 10) with samtools. MACS2 called ≈61k FAIRE peaks. The STARR peaks overlapped with the FAIRE peaks with precisions of 52.7% (high-confidence peaks), 40.6% (medium-confidence peaks), and 22.5% (low-confidence peaks).