Human pol II promoter prediction: time series descriptors and machine learning

Human pol II promoter prediction: time series descriptors and machine learning
复制标题

DOI:
10.1093/nar/gki271
复制
发表时间:
2005-01-01
影响因子:
14.9
通讯作者:
Sharma, P
Sharma, P
中科院分区:
生物学2区
文献类型:
--
作者:
Gangal, R;Sharma, P

文献摘要

被引文献

相似文献

尽管迄今为止已经开发了几种计算机启动子预测方法,但它们在预测性能方面仍然有限。这些限制是由于选择启动子的适当特征以将其与非启动子区分开的挑战以及机器学习算法的泛化或预测能力。在本文中,我们试图定义一种新的方法,通过使用独特的描述符和机器学习方法识别真核生物聚合酶II启动子。在这项研究中,非线性时间序列描述符沿着与非线性机器学习算法,如支持向量机(SVM),用于区分启动子和非启动子区域。这里的基本思想是使用不依赖于一级DNA序列的描述符,并提供启动子和非启动子区域之间的明确区分。建立在一组1000个启动子和1500个非启动子序列上的分类模型显示出87%的10倍交叉验证准确率,并且独立测试集在启动子和非启动子鉴定中的准确率> 85%。该方法正确鉴定了人类22号染色体的所有20个实验验证的启动子。高灵敏度和选择性表明,n-mer频率沿着非线性时间序列描述符,如李雅普诺夫分量稳定性和Tsallis熵,以及监督机器学习方法,如SVM,可用于鉴定pol II启动子。
Although several in silico promoter prediction methods have been developed to date, they are still limited in predictive performance. The limitations are due to the challenge of selecting appropriate features of promoters that distinguish them from non-promoters and the generalization or predictive ability of the machine-learning algorithms. In this paper we attempt to define a novel approach by using unique descriptors and machine-learning methods for the recognition of eukaryotic polymerase II promoters. In this study, non-linear time series descriptors along with non-linear machine-learning algorithms, such as support vector machine (SVM), are used to discriminate between promoter and non-promoter regions. The basic idea here is to use descriptors that do not depend on the primary DNA sequence and provide a clear distinction between promoter and non-promoter regions. The classification model built on a set of 1000 promoter and 1500 non-promoter sequences, showed a 10-fold cross-validation accuracy of 87% and an independent test set had an accuracy > 85% in both promoter and non-promoter identification. This approach correctly identified all 20 experimentally verified promoters of human chromosome 22. The high sensitivity and selectivity indicates that n-mer frequencies along with non-linear time series descriptors, such as Lyapunov component stability and Tsallis entropy, and supervised machine-learning methods, such as SVMs, can be useful in the identification of pol II promoters.