Predicting genetic regulatory response using classification

Predicting genetic regulatory response using classification
复制标题

DOI:
10.1093/bioinformatics/bth923
复制
发表时间:
2004-08-04
期刊:
影响因子:
5.8
通讯作者:
Leslie, Christina
Leslie, Christina
中科院分区:
生物学3区
文献类型:
--
作者:
Middendorf, Manuel;Kundaje, Anshul;Leslie, Christina

文献摘要

被引文献

相似文献

动机:通过分析高通量基因组数据来研究简单模式生物中的基因调控机制已经成为计算生物学的一个中心问题。文献中的大多数方法都集中在寻找一些强有力的监管模式或从训练数据中学习描述性模型。然而,这些方法还不足以准确预测哪些基因将在新的或保留的实验中上调或下调。通过引入一个预测方法,这个问题,我们可以使用强大的工具,从机器学习和评估我们的predictions.Results的统计意义:我们提出了一种新的基于分类的学习方法来预测基因调控反应。我们的方法的动机是假设,在简单的生物体,如酿酒酵母,我们可以学习一个决策规则,用于预测是否一个基因是上调或下调在一个特定的实验的基础上(1)结合位点序列(“基序”)在基因的调控区的存在和(2)的表达水平的调节因子,如转录因子在实验中(“父母”)。因此,我们的学习任务整合了两个性质不同的数据源:跨多个扰动和突变实验的全基因组cDNA微阵列数据沿着调控序列的基序分布数据。我们将预测实值基因表达测量的回归任务转换为预测+1和-1标签的分类任务,对应于超出微阵列测量中生物和测量噪声水平的上调和下调。所采用的学习算法是基于边缘的决策树泛化,交替决策树。这个大的利润分类器是足够灵活的,允许复杂的逻辑功能,但足够简单,让洞察基因调控的组合机制。我们观察到令人鼓舞的预测精度的实验的基础上的Gasch S。酿酒酵母数据集,我们表明,我们可以准确地预测上调和下调举行了实验。我们还展示了如何从各种压力反应的学习模型中提取重要的调节器,图案和图案-调节器对。因此,我们的方法提供了预测假设,建议生物实验,并提供了可解释的遗传调控网络的结构的见解。
Motivation: Studying gene regulatory mechanisms in simple model organisms through analysis of high-throughput genomic data has emerged as a central problem in computational biology. Most approaches in the literature have focused either on finding a few strong regulatory patterns or on learning descriptive models from training data. However, these approaches are not yet adequate for making accurate predictions about which genes will be up-or down-regulated in new or held-out experiments. By introducing a predictive methodology for this problem, we can use powerful tools from machine learning and assess the statistical significance of our predictions.Results: We present a novel classification-based method for learning to predict gene regulatory response. Our approach is motivated by the hypothesis that in simple organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up-or down-regulated in a particular experiment based on (1) the presence of binding site subsequences ('motifs') in the gene's regulatory region and (2) the expression levels of regulators such as transcription factors in the experiment ('parents'). Thus, our learning task integrates two qualitatively different data sources: genome-wide cDNA microarray data across multiple perturbation and mutant experiments along with motif profile data from regulatory sequences. We convert the regression task of predicting real-valued gene expression measurements to a classification task of predicting +1 and -1 labels, corresponding to up-and down-regulation beyond the levels of biological and measurement noise in microarray measurements. The learning algorithm employed is boosting with a margin-based generalization of decision trees, alternating decision trees. This large-margin classifier is sufficiently flexible to allow complex logical functions, yet sufficiently simple to give insight into the combinatorial mechanisms of gene regulation. We observe encouraging prediction accuracy on experiments based on the Gasch S. cerevisiae dataset, and we show that we can accurately predict up-and down-regulation on held-out experiments. We also show how to extract significant regulators, motifs and motif-regulator pairs from the learned models for various stress responses. Our method thus provides predictive hypotheses, suggests biological experiments, and provides interpretable insight into the structure of genetic regulatory networks.