Extracting relations between promoter sequences and their strengths from microarray data

Extracting relations between promoter sequences and their strengths from microarray data
复制标题

DOI:
10.1093/bioinformatics/bti094
复制
发表时间:
2005-04-01
期刊:
影响因子:
5.8
通讯作者:
Asai, K
Asai, K
中科院分区:
生物学3区
文献类型:
--
作者:
Kiryu, H;Oshima, T;Asai, K

文献摘要

被引文献

相似文献

动机:启动子序列与其强度之间的关系在20世纪80年代得到了广泛的研究。尽管这些研究发现了很强的序列-强度相关性,但他们复杂的实验方法的成本太高,无法应用于大量的启动子。相反,最近微阵列数据的增加使我们能够将数千个基因的表达与其DNA序列进行比较。结果:我们利用大肠杆菌微阵列数据研究了启动子序列与其强度之间的关系。我们使用一个简单的权重矩阵对这些关系进行建模,并使用一种新的支持向量回归方法对其进行优化。观察到启动子序列‘-35’和‘-10’区域的几个非共识碱基对启动子强度起正作用,而某些共识碱基对启动子强度的影响较小。我们分析了观察到的基因表达偏离启动子强度预测的异常值,并确定了几个由于多个启动子而增强表达的基因和受转录因子强烈调控的基因。我们的方法适用于启动子序列和微阵列数据均可获得的其他原核生物。
Motivation: The relations between the promoter sequences and their strengths were extensively studied in the 1980s. Although these studies uncovered strong sequence-strength correlations, the cost of their elaborate experimental methods have been too high to be applied to a large number of promoters. On the contrary, a recent increase in the microarray data allows us to compare thousands of gene expressions with their DNA sequences.Results: We studied the relations between the promoter sequences and their strengths using the Escherichia coli microarray data. We modeled those relations using a simple weight matrix, which was optimized with a novel support vector regression method. It was observed that several non-consensus bases in the '-35' and '-10' regions of promoter sequences act positively on the promoter strength and that certain consensus bases have a minor effect on the strength. We analyzed outliers for which the observed gene expressions deviate from the promoter strength predictions, and identified several genes with enhanced expressions due to multiple promoters and genes under strong regulation by transcription factors. Our method is applicable to other procaryotes for which both the promoter sequences and the microarray data are available.