F-Seq2: improving the feature density based peak caller with dynamic statistics.

F-Seq2: improving the feature density based peak caller with dynamic statistics.
复制标题

F-Seq 2:利用动态统计改进基于特征密度的峰值调用者。

DOI:
10.1093/nargab/lqab012
复制
发表时间:
2021-03
影响因子:
4.6
通讯作者:
Boyle AP
Boyle AP
中科院分区:
其他
文献类型:
--
作者:
Zhao N;Boyle AP

文献摘要

参考文献

被引文献

相似文献

通过使用高通量测序(HTS)技术在全基因组水平上捕获基因组和表观基因组特征。通过将观察到的读数分布与随机预期进行比较,峰识别描绘了HTS实验中鉴定的特征,例如开放的染色质区域和转录因子结合位点。自其引入以来,F-Seq已被广泛使用,并被证明是DNase I超敏位点(DNase-seq)数据的最灵敏和最准确的峰调用者。然而,第一个版本(F-Seq 1)有两个关键的限制:缺乏对用户输入控制数据集的支持,以及测试统计报告不佳。这些限制了其捕获峰预测中背景分布固有的系统和实验偏差的能力,以及随后通过置信度对预测峰进行排名的能力。为了解决这些局限性,我们提出了F-Seq 2,它结合了核密度估计和动态“连续”泊松测试来考虑局部偏差并准确地对候选峰进行排名。F-Seq 2的输出适用于不可重现的发现率分析,因为测试统计量是针对单个候选峰计算的,允许直接比较重复的预测。这些改进显著提高了F-Seq 2在ATAC-seq和ChIP-seq数据集上的性能,在精确度和召回率方面优于ENCODE联盟使用的竞争峰值调用器。
Genomic and epigenomic features are captured at a genome-wide level by using high-throughput sequencing (HTS) technologies. Peak calling delineates features identified in HTS experiments, such as open chromatin regions and transcription factor binding sites, by comparing the observed read distributions to a random expectation. Since its introduction, F-Seq has been widely used and shown to be the most sensitive and accurate peak caller for DNase I hypersensitive site (DNase-seq) data. However, the first release (F-Seq1) has two key limitations: lack of support for user-input control datasets, and poor test statistic reporting. These constrain its ability to capture systematic and experimental biases inherent to the background distributions in peak prediction, and to subsequently rank predicted peaks by confidence. To address these limitations, we present F-Seq2, which combines kernel density estimation and a dynamic ‘continuous’ Poisson test to account for local biases and accurately rank candidate peaks. The output of F-Seq2 is suitable for irreproducible discovery rate analysis as test statistics are calculated for individual candidate summits, allowing direct comparison of predictions across replicates. These improvements significantly boost the performance of F-Seq2 for ATAC-seq and ChIP-seq datasets, outperforming competing peak callers used by the ENCODE Consortium in terms of precision and recall.
DOI: 10.1214/aoms/1177704472
发表时间: 1962-01-01
影响因子: --
作者:
PARZEN, E
通讯作者: PARZEN, E
DOI: 10.1186/1748-7188-2-15
发表时间: 2007-12-11
影响因子: 1
作者:
Touzet, Helene;Varre, Jean-Stephane
通讯作者: Varre, Jean-Stephane
DOI: 10.1371/journal.pcbi.1002638
发表时间: 2012
影响因子: 4.3
作者:
Guo Y;Mahony S;Gifford DK
通讯作者: Gifford DK
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1186/gb-2009-10-3-r25
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Langmead B;Trapnell C;Pop M;Salzberg SL
通讯作者: Salzberg SL