ChIPulate: A comprehensive ChIP-seq simulation pipeline

ChIPulate: A comprehensive ChIP-seq simulation pipeline
复制标题

DOI:
10.1371/journal.pcbi.1006921
复制
发表时间:
2019-03-01
影响因子:
4.3
通讯作者:
Siddharthan, Rahul
Siddharthan, Rahul
中科院分区:
生物学2区
文献类型:
--
作者:
Datta, Vishaka;Hannenhalli, Sridhar;Siddharthan, Rahul

文献摘要

被引文献

相似文献

ChIP-seq(染色质免疫沉淀,随后测序)是一种高通量技术,用于鉴定在体内被特定蛋白质结合的基因组区域,例如,转录因子(TF)。已知生物因素,如染色质状态、间接和合作结合,以及实验因素,如抗体质量、交联和PCR偏倚,会影响ChIP-seq实验的结果。然而,这些因素对ChIP-seq数据推断的相对影响尚不完全清楚。在这里,通过详细的ChIP-seq模拟管道ChIPulate,我们评估了各种生物和实验变异来源对ChIP-seq实验的几个结果的影响,即,TF结合基序的可恢复性、TF-DNA结合检测的准确性、推断的TF-DNA结合强度的灵敏度和可靠地推断结合强度所需的重复次数。我们发现,TF基序可以回收,尽管穷人和不均匀的提取和PCR扩增效率。然而,基序的恢复在更大程度上受到协同或间接结合的位点的分数的影响。重要的是,我们的模拟显示,准确测量高亲和力位点的体内占用所需的ChIP-seq重复次数大于推荐的社区标准。我们的研究结果建立了从ChIP-seq推断蛋白质-DNA结合的准确性的统计限制,并表明增加平均提取效率而不是扩增效率将更好地提高灵敏度。运行ChIPulate的源代码和说明可以在https://github.com/vishakad/chipulate.Author上找到。摘要DNA结合蛋白在生物学中起着许多关键作用,如基因表达的转录调控和染色质修饰。ChIP-seq(染色质免疫沉淀,然后进行高通量测序)是一种广泛使用的实验技术,用于在细胞内全基因组范围内识别特定目标蛋白质的DNA结合位点。使用特异性抗体选择性地提取来自被感兴趣的蛋白质(通常是转录因子(TF))结合的基因组区域的DNA片段,使用PCR扩增,并测序。将序列映射到参考基因组。其中许多序列映射的区域(称为峰)用于推断TF结合基因座(峰)的位置、在这些基因座的体内占有率以及TF对其显示结合亲和力的序列模式(基序)。但是TF占用和基序推断容易受到几种生物学和实验变异来源的影响,这些变异来源知之甚少,难以直接评估。在这里,我们模拟ChIP-seq协议的关键步骤,目的是估计各种变异来源对基序推断和结合亲和力估计的相对影响。除了提供具体的见解和建议外,我们还提供了一个通用框架来模拟ChIP-seq实验中的序列读取,这将大大有助于开发旨在分析ChIP-seq数据的软件。
ChIP-seq (Chromatin Immunoprecipitation followed by sequencing) is a high-throughput technique to identify genomic regions that are bound in vivo by a particular protein, e.g., a transcription factor (TF). Biological factors, such as chromatin state, indirect and cooperative binding, as well as experimental factors, such as antibody quality, cross-linking, and PCR biases, are known to affect the outcome of ChIP-seq experiments. However, the relative impact of these factors on inferences made from ChIP-seq data is not entirely clear. Here, via a detailed ChIP-seq simulation pipeline, ChIPulate, we assess the impact of various biological and experimental sources of variation on several outcomes of a ChIP-seq experiment, viz., the recoverability of the TF binding motif, accuracy of TF-DNA binding detection, the sensitivity of inferred TF-DNA binding strength, and number of replicates needed to confidently infer binding strength. We find that the TF motif can be recovered despite poor and non-uniform extraction and PCR amplification efficiencies. The recovery of the motif is, however, affected to a larger extent by the fraction of sites that are either cooperatively or indirectly bound. Importantly, our simulations reveal that the number of ChIP-seq replicates needed to accurately measure in vivo occupancy at high-affinity sites is larger than the recommended community standards. Our results establish statistical limits on the accuracy of inferences of protein-DNA binding from ChIP-seq and suggest that increasing the mean extraction efficiency, rather than amplification efficiency, would better improve sensitivity. The source code and instructions for running ChIPulate can be found at https://github.com/vishakad/chipulate.Author summary DNA-binding proteins perform many key roles in biology, such as transcriptional regulation of gene expression and chromatin modification. ChIP-seq (Chromatin immunoprecipitation followed by high-throughput sequencing) is a widely used experimental technique to identify DNA-binding sites of specific proteins of interest, within cells, genome-wide. DNA fragments from genomic regions that are bound by a protein of interest, often a transcription factor (TF), are selectively extracted using specific antibodies, amplified using PCR, and sequenced. The sequences are mapped to the reference genome. Regions where many sequences map, called peaks, are used to infer the location of TF-bound loci (peaks), in vivo occupancy at those loci, and the sequence pattern (motif) to which the TF shows a binding affinity. But TF occupancy and motif inference are vulnerable to several biological and experimental sources of variation that are poorly understood and difficult to assess directly. Here, we simulate key steps of the ChIP-seq protocol with the aim of estimating the relative effects of various sources of variations on motif inference and binding affinity estimations. Besides providing specific insights and recommendations, we provide a general framework to simulate sequence reads in a ChIP-seq experiment, which should considerably aid in the development of software aimed at analyzing ChIP-seq data.