DNA sequence+shape kernel enables alignment-free modeling of transcription factor binding.

DNA sequence+shape kernel enables alignment-free modeling of transcription factor binding.
复制标题

DOI:
10.1093/bioinformatics/btx336
复制
发表时间:
2017-10-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Noble WS
Noble WS
中科院分区:
其他
文献类型:
--
作者:
Ma W;Yang L;Rohs R;Noble WS

文献摘要

参考文献

被引文献

相似文献

转录因子 (TF) 与特定的 DNA 序列基序结合。多项证据表明,TF-DNA 结合部分由局部 DNA 形状的特性介导:小沟的宽度、相邻碱基对的相对方向等。已经开发了几种方法来共同考虑 DNA 序列和形状特性,以预测 TF 结合亲和力。然而,这些方法的局限性在于它们通常需要一组对齐的 TF 结合位点的训练集。我们描述了一个序列“+”形状内核,它利用 DNA 序列和形状信息来更好地理解蛋白质-DNA 结合偏好和亲和力。该内核扩展了现有的基于 k-mer 的序列内核,基于最近描述的二不匹配内核。使用源自通用蛋白质结合微阵列 (uPBM)、基因组背景 PBM (gcPBM) 和 SELEX-seq 数据的三个体外基准数据集,我们证明整合 DNA 形状信息可以提高我们预测蛋白质-DNA 结合亲和力的能力。特别是,我们观察到 (i) k-spectrum + shape 模型比经典的 k-spectrum kernel 表现更好,特别是对于较小的 k 值; (ii) 对于较大的 k,di-mismatch 内核的性能优于 k-mer 内核; (iii) 对于中间 k 值,di-mismatch + shape 内核的性能优于 di-mismatch 内核。该软件可从 https://bitbucket.org/wenxiu/sequence-shape.git 获取。 补充数据可在生物信息学在线获取。
Transcription factors (TFs) bind to specific DNA sequence motifs. Several lines of evidence suggest that TF-DNA binding is mediated in part by properties of the local DNA shape: the width of the minor groove, the relative orientations of adjacent base pairs, etc. Several methods have been developed to jointly account for DNA sequence and shape properties in predicting TF binding affinity. However, a limitation of these methods is that they typically require a training set of aligned TF binding sites. We describe a sequence + shape kernel that leverages DNA sequence and shape information to better understand protein-DNA binding preference and affinity. This kernel extends an existing class of k-mer based sequence kernels, based on the recently described di-mismatch kernel. Using three in vitro benchmark datasets, derived from universal protein binding microarrays (uPBMs), genomic context PBMs (gcPBMs) and SELEX-seq data, we demonstrate that incorporating DNA shape information improves our ability to predict protein-DNA binding affinity. In particular, we observe that (i) the k-spectrum + shape model performs better than the classical k-spectrum kernel, particularly for small k values; (ii) the di-mismatch kernel performs better than the k-mer kernel, for larger k; and (iii) the di-mismatch + shape kernel performs better than the di-mismatch kernel for intermediate k values. The software is available at https://bitbucket.org/wenxiu/sequence-shape.git. Supplementary data are available at Bioinformatics online.
DOI: 10.1016/j.cell.2011.10.053
发表时间: 2011-12-09
期刊: Cell
影响因子: 64.5
作者:
Slattery M;Riley T;Liu P;Abe N;Gomez-Alcala P;Dror I;Zhou T;Rohs R;Honig B;Bussemaker HJ;Mann RS
通讯作者: Mann RS
DOI: 10.1371/journal.pcbi.1000916
发表时间: 2010-09-01
影响因子: 4.3
作者:
Agius, Phaedra;Arvey, Aaron;Leslie, Christina
通讯作者: Leslie, Christina
DOI: 10.1101/gr.184671.114
发表时间: 2015-09
期刊: Genome research
影响因子: 7
作者:
Dror I;Golan T;Levy C;Rohs R;Mandel-Gutfreund Y
通讯作者: Mandel-Gutfreund Y
DOI: 10.1038/nbt1246
发表时间: 2006-11-01
影响因子: 46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.
DOI: 10.1093/bioinformatics/btv735
发表时间: 2016-04-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Chiu TP;Comoglio F;Zhou T;Yang L;Paro R;Rohs R
通讯作者: Rohs R