Prediction accuracy of regulatory elements from sequence varies by functional sequencing technique.

Prediction accuracy of regulatory elements from sequence varies by functional sequencing technique.
复制标题

DOI:
10.3389/fcimb.2023.1182567
复制
发表时间:
2023
影响因子:
5.7
通讯作者:
--
中科院分区:
医学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

各种基于测序的方法用于以全基因组方式鉴定和表征顺式调控元件的活性。这些技术中的一些依赖于间接标记物,如组蛋白修饰(ChIP-seq与组蛋白抗体)或染色质可及性(ATAC-seq、DNase-seq、FAIRE-seq),而其他技术使用直接测量,如测量DNA序列的增强子特性的附加型测定(STARR-seq)和直接测量转录因子的结合(ChIP-seq与转录因子特异性抗体)。顺式调控元件如增强子、启动子和阻遏物的活性由它们的序列和次级过程如染色质可及性、DNA甲基化和结合的组蛋白标记物决定。这里,采用机器学习模型来评估通过各种常用测序技术鉴定的顺式调节元件可以单独通过其潜在序列预测的准确性,以区分反映序列内容的顺式调节活性与次级过程。模型在D.通过DNase-seq和STARR-seq鉴定的黑腹鱼序列比在通过H3 K4 me 1、H3 K4 me 3和H3 K27 ac ChIP-seq、FAIRE-seq和ATAC-seq鉴定的序列上训练的模型显著更准确。这些结果表明,通过DNase-seq和STARR-seq检测到的活性可以在很大程度上通过潜在的DNA序列来解释,而不依赖于次级过程。在实验上,使用荧光素酶测定法测试了DNase-seq和H3 K4 me 1 ChIP-seq序列的子集的增强子活性,并与先前对STARR-seq序列进行的测试进行了比较。实验数据表明,STARR-seq序列基本上富集了增强子特异性活性,而DNase-seq和H3 K4 me 1 ChIP-seq序列则没有。综上所述,这些结果表明,DNase-seq方法鉴定了广泛的一类调控元件,其中增强子是其子集,并且相关数据适合于训练用于单独从序列检测调控活性的模型,STARR-seq数据最适合于训练增强子特异性序列模型,和H3 K4 me 1 ChIP-seq数据不太适合于训练和评估用于顺式调控元件预测的基于序列的模型。
Various sequencing based approaches are used to identify and characterize the activities of cis-regulatory elements in a genome-wide fashion. Some of these techniques rely on indirect markers such as histone modifications (ChIP-seq with histone antibodies) or chromatin accessibility (ATAC-seq, DNase-seq, FAIRE-seq), while other techniques use direct measures such as episomal assays measuring the enhancer properties of DNA sequences (STARR-seq) and direct measurement of the binding of transcription factors (ChIP-seq with transcription factor-specific antibodies). The activities of cis-regulatory elements such as enhancers, promoters, and repressors are determined by their sequence and secondary processes such as chromatin accessibility, DNA methylation, and bound histone markers. Here, machine learning models are employed to evaluate the accuracy with which cis-regulatory elements identified by various commonly used sequencing techniques can be predicted by their underlying sequence alone to distinguish between cis-regulatory activity that is reflective of sequence content versus secondary processes. Models trained and evaluated on D. melanogaster sequences identified through DNase-seq and STARR-seq are significantly more accurate than models trained on sequences identified by H3K4me1, H3K4me3, and H3K27ac ChIP-seq, FAIRE-seq, and ATAC-seq. These results suggest that the activity detected by DNase-seq and STARR-seq can be largely explained by underlying DNA sequence, independent of secondary processes. Experimentally, a subset of DNase-seq and H3K4me1 ChIP-seq sequences were tested for enhancer activity using luciferase assays and compared with previous tests performed on STARR-seq sequences. The experimental data indicated that STARR-seq sequences are substantially enriched for enhancer-specific activity, while the DNase-seq and H3K4me1 ChIP-seq sequences are not. Taken together, these results indicate that the DNase-seq approach identifies a broad class of regulatory elements of which enhancers are a subset and the associated data are appropriate for training models for detecting regulatory activity from sequence alone, STARR-seq data are best for training enhancer-specific sequence models, and H3K4me1 ChIP-seq data are not well suited for training and evaluating sequence-based models for cis-regulatory element prediction.
DOI: 10.1002/stem.1871
发表时间: 2015-02
期刊: STEM CELLS
影响因子: 5.2
作者:
Murtha, Matthew;Strino, Francesco;Tokcaer-Keskin, Zeynep;Bayin, N. Sumru;Shalabi, Doaa;Xi, Xiangmei;Kluger, Yuval;Dailey, Lisa
通讯作者: Dailey, Lisa