Enhanced regulatory sequence prediction using gapped k-mer features.

Enhanced regulatory sequence prediction using gapped k-mer features.
复制标题

DOI:
10.1371/journal.pcbi.1003711
复制
发表时间:
2014-07
影响因子:
4.3
通讯作者:
Beer MA
Beer MA
中科院分区:
生物学2区
文献类型:
--
作者:
Ghandi M;Lee D;Mohammad-Noori M;Beer MA

文献摘要

被引文献

相似文献

长度为\(k\)的寡聚体,或\(k\)-聚体,是对DNA和蛋白质序列的性质及功能进行建模时方便且广泛使用的特征。然而,\(k\)-聚体存在固有局限性,即如果增加参数\(k\)以解析更长的特征,观察到任何特定\(k\)-聚体的概率会变得非常小,并且\(k\)-聚体计数接近一个二元变量,大多数\(k\)-聚体不存在,只有少数出现一次。因此,一旦\(k\)变大,任何使用\(k\)-聚体作为特征的统计学习方法都会容易受到有噪声的训练集\(k\)-聚体频率的影响。为解决这个问题,我们引入了使用间隔\(k\)-聚体的替代特征集、一种新的分类器\(gkm - SVM\)以及一种对\(k\)-聚体频率进行稳健估计的通用方法。为使该方法适用于大规模全基因组应用,我们开发了一种用于计算核矩阵的高效树形数据结构。我们表明,与我们原来的\(kmer - SVM\)和其他替代方法相比,我们的\(gkm - SVM\)在预测功能基因组调控元件和组织特异性增强子方面具有显著提高的准确性,精度提高了多达两倍。然后我们表明,\(gkm - SVM\)在人类ENCODE ChIP - seq数据集上始终优于\(kmer - SVM\),并且使用朴素贝叶斯分类器进一步证明了我们方法的通用性。尽管是为调控序列分析而开发的,但这些方法可应用于任何序列分类问题。 基因组调控元件(增强子、启动子和绝缘子)控制其靶基因的表达,并且人们普遍认为它们通过改变蛋白质浓度在人类发育和疾病中起关键作用。理解增强子的一个基本步骤是开发基于DNA序列的模型来预测调控元件的组织特异性活性。此类模型既有助于识别通过直接转录因子结合影响增强子活性的分子途径,也有助于直接评估特定常见或罕见遗传变异对增强子功能的影响。我们之前使用\(k\)-聚体支持向量机(\(kmer - SVM\))开发了一个成功的基于序列的增强子预测模型。在这里,我们解决了\(kmer - SVM\)方法的一个重大局限性,并提出了一种使用间隔\(k\)-聚体(\(gkm - SVM\))的替代方法,该方法在所有测试案例中都显示出显著提高的准确性。虽然我们专注于增强子和转录因子结合,但我们的方法可应用于改善更广泛的一类序列分析问题,包括蛋白质和RNA。此外,我们预计大多数基于\(k\)-聚体的方法只需使用我们在本文中提出的广义\(k\)-聚体计数方法就可以得到显著改进。我们相信这个改进的模型将对我们理解人类调控系统做出重大贡献。
Oligomers of length k, or k-mers, are convenient and widely used features for modeling the properties and functions of DNA and protein sequences. However, k-mers suffer from the inherent limitation that if the parameter k is increased to resolve longer features, the probability of observing any specific k-mer becomes very small, and k-mer counts approach a binary variable, with most k-mers absent and a few present once. Thus, any statistical learning approach using k-mers as features becomes susceptible to noisy training set k-mer frequencies once k becomes large. To address this problem, we introduce alternative feature sets using gapped k-mers, a new classifier, gkm-SVM, and a general method for robust estimation of k-mer frequencies. To make the method applicable to large-scale genome wide applications, we develop an efficient tree data structure for computing the kernel matrix. We show that compared to our original kmer-SVM and alternative approaches, our gkm-SVM predicts functional genomic regulatory elements and tissue specific enhancers with significantly improved accuracy, increasing the precision by up to a factor of two. We then show that gkm-SVM consistently outperforms kmer-SVM on human ENCODE ChIP-seq datasets, and further demonstrate the general utility of our method using a Naïve-Bayes classifier. Although developed for regulatory sequence analysis, these methods can be applied to any sequence classification problem. Genomic regulatory elements (enhancers, promoters, and insulators) control the expression of their target genes and are widely believed to play a key role in human development and disease by altering protein concentrations. A fundamental step in understanding enhancers is the development of DNA sequence-based models to predict the tissue specific activity of regulatory elements. Such models facilitate both the identification of the molecular pathways which impinge on enhancer activity through direct transcription factor binding, and the direct evaluation of the impact of specific common or rare genetic variants on enhancer function. We have previously developed a successful sequence-based model for enhancer prediction using a k-mer support vector machine (kmer-SVM). Here, we address a significant limitation of the kmer-SVM approach and present an alternative method using gapped k-mers (gkm-SVM) which exhibits dramatically improved accuracy in all test cases. While we focus on enhancers and transcription factor binding, our method can be applied to improve a much broader class of sequence analysis problems, including proteins and RNA. In addition, we expect that most k-mer based methods can be significantly improved by simply using the generalized k-mer count method that we present in this paper. We believe this improved model will enable significant contributions to our understanding of the human regulatory system.