A novel k-mer set memory (KSM) motif representation improves regulatory variant prediction.

A novel k-mer set memory (KSM) motif representation improves regulatory variant prediction.
复制标题

DOI:
10.1101/gr.226852.117
复制
发表时间:
2018-06
期刊:
影响因子:
7
通讯作者:
Gifford DK
Gifford DK
中科院分区:
生物学1区
文献类型:
--
作者:
Guo Y;Tian K;Zeng H;Guo X;Gifford DK

文献摘要

参考文献

被引文献

相似文献

转录因子(TF)序列结合特异性的表征和发现对于理解基因调控网络和解释疾病相关非编码遗传变异的影响至关重要。我们提出了一种新的TF结合基序表示,k-mer集记忆(KSM),它由一组对齐的k-mer,在TF结合位点过度,和一种新的方法称为KMAC从头发现KSM。我们发现KSM比位置权重矩阵(PWM)模型和其他更复杂的基序模型更准确地预测体内结合位点。此外,KSM在预测体外结合位点方面优于PWM和更复杂的基序模型。KMAC还在比五种最先进的基序发现方法更多的实验中识别正确的基序。此外,KSM衍生的特征在预测表达数量性状基因座(eQTL)等位基因的差异调节活性方面优于PWM和深度学习模型衍生的序列特征。最后,我们将KMAC应用于1600个ENCODE TF ChIP-seq数据集,并创建了KSM和PWM基序的公共资源。我们希望KSM表示和KMAC方法将是有价值的TF结合特异性的特征和解释非编码遗传变异的影响。
The representation and discovery of transcription factor (TF) sequence binding specificities is critical for understanding gene regulatory networks and interpreting the impact of disease-associated noncoding genetic variants. We present a novel TF binding motif representation, the k-mer set memory (KSM), which consists of a set of aligned k-mers that are overrepresented at TF binding sites, and a new method called KMAC for de novo discovery of KSMs. We find that KSMs more accurately predict in vivo binding sites than position weight matrix (PWM) models and other more complex motif models across a large set of ChIP-seq experiments. Furthermore, KSMs outperform PWMs and more complex motif models in predicting in vitro binding sites. KMAC also identifies correct motifs in more experiments than five state-of-the-art motif discovery methods. In addition, KSM-derived features outperform both PWM and deep learning model derived sequence features in predicting differential regulatory activities of expression quantitative trait loci (eQTL) alleles. Finally, we have applied KMAC to 1600 ENCODE TF ChIP-seq data sets and created a public resource of KSM and PWM motifs. We expect that the KSM representation and KMAC method will be valuable in characterizing TF binding specificities and in interpreting the effects of noncoding genetic variations.
DOI: 10.1126/science.1162327
发表时间: 2009-06-26
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Badis G;Berger MF;Philippakis AA;Talukder S;Gehrke AR;Jaeger SA;Chan ET;Metzler G;Vedenko A;Chen X;Kuznetsov H;Wang CF;Coburn D;Newburger DE;Morris Q;Hughes TR;Bulyk ML
通讯作者: Bulyk ML
DOI: 10.1093/bioinformatics/btl243
发表时间: 2006-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fratkin, Eugene;Naughton, Brian T.;Batzoglou, Serafim
通讯作者: Batzoglou, Serafim
DOI: 10.1038/nbt1246
发表时间: 2006-11-01
影响因子: 46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.
DOI: 10.1093/nar/30.5.1255
发表时间: 2002-03-01
影响因子: 14.9
作者:
Bulyk, ML;Johnson, PLF;Church, GM
通讯作者: Church, GM
DOI: 10.1002/humu.23197
发表时间: 2017-09
期刊: Human mutation
影响因子: 3.9
作者:
Kreimer A;Zeng H;Edwards MD;Guo Y;Tian K;Shin S;Welch R;Wainberg M;Mohan R;Sinnott-Armstrong NA;Li Y;Eraslan G;Amin TB;Tewhey R;Sabeti PC;Goke J;Mueller NS;Kellis M;Kundaje A;Beer MA;Keles S;Gifford DK;Yosef N
通讯作者: Yosef N