ProSampler: an ultrafast and accurate motif finder in large ChIP-seq datasets for combinatory motif discovery

ProSampler: an ultrafast and accurate motif finder in large ChIP-seq datasets for combinatory motif discovery
复制标题

ProSampler:大型 ChIP-seq 数据集中的超快且准确的基序查找器,用于发现组合基序

DOI:
10.1093/bioinformatics/btz290
复制
发表时间:
2019-11-15
期刊:
影响因子:
5.8
通讯作者:
Su, Zhengchang
Su, Zhengchang
中科院分区:
生物学3区
文献类型:
--
作者:
Li, Yang;Ni, Pengyu;Su, Zhengchang

文献摘要

被引文献

相似文献

动机 大量转录因子(TF)的ChIP-seq数据集的可用性为识别基因组中的所有TF结合位点提供了前所未有的机会。然而,由于缺乏一种高效、准确的工具,不仅可以在非常大的数据集中找到目标基序,还可以找到合作基序,这一进展受到了阻碍。 结果 本文提出了一种基于新的计数方法和Gibbs采样器的超快速准确的基序发现算法ProSampler。ProSampler的运行速度比现有最快的工具快几个数量级,同时通常更准确地识别目标TF和合作者的基序。因此,ProSampler可以极大地促进鉴定基因组中的整个顺式调控密码的努力。 可用性 源代码和二进制文件可以在https://github.com/zhengchangsulab/prosampler上免费下载。它在C++中实现,并在Linux、macOS和MS Windows平台上得到支持。 补充资料 补充材料可在生物信息学在线。
MOTIVATION The availability of numerous ChIP-seq datasets for transcription factors (TF) has provided an unprecedented opportunity to identify all TF binding sites in genomes. However, the progress has been hindered by the lack of a highly efficient and accurate tool to find not only the target motifs, but also cooperative motifs in very big datasets. RESULTS We herein present an ultra-fast and accurate motif-finding algorithm, ProSampler, based on a novel numeration method and Gibbs sampler. ProSampler runs orders of magnitude faster than the fastest existing tools while often more accurately identifying motifs of both the target TFs and cooperators. Thus, ProSampler can greatly facilitate the efforts to identify the entire cis-regulatory code in genomes. AVAILABILITY Source code and binaries are freely available for download at https://github.com/zhengchangsulab/prosampler. It was implemented in C ++ and supported on Linux, macOS and MS Windows platforms. SUPPLEMENTARY INFORMATION Supplementary materials are available at Bioinformatics online.