Learning gene regulatory networks from only positive and unlabeled data.

Learning gene regulatory networks from only positive and unlabeled data.
复制标题

DOI:
10.1186/1471-2105-11-228
复制
发表时间:
2010-05-05
期刊:
影响因子:
3
通讯作者:
Ceccarelli M
Ceccarelli M
中科院分区:
生物学4区
文献类型:
--
作者:
Cerulo L;Elkan C;Ceccarelli M

文献摘要

参考文献

被引文献

相似文献

最近,已利用监督的学习方法从基因表达数据中重建基因调节网络。网络的重建被建模为每对基因的二进制分类问题。训练了统计分类器,以识别基因对的激活曲线之间的关系。事实证明,这种方法的表现要优于以前的无监督方法。但是,监督的方法提出了空旷的问题。特别是,尽管可以安全地认为已知的调节连接是积极的训练示例,但获得负面示例并不简单,因为通常无法获得确定的知识,即给定的基因不相互作用。 对数据挖掘研究的最新进展是一种能够从仅积极和未标记的示例中学习分类器的方法,而这些示例不需要标记为负面示例。应用于基因调节网络的重建,我们表明该方法显着优于机器学习方法的当前状态。我们使用模拟和实验数据评估新方法,并获得重大的性能改进。 与基因网络推断的无监督方法相比,有监督的方法可能更准确,但是对于训练,他们需要一组完整的已知调节连接。本文提出的一种可以仅使用正面和未标记数据培训的有监督方法对于推断基因调节网络的任务尤其有益,因为只有一组不完整的已知调节连接集可在RegulondB等公共数据库中获得,例如TRRD,KEGG,TransFac和IPA。
Recently, supervised learning methods have been exploited to reconstruct gene regulatory networks from gene expression data. The reconstruction of a network is modeled as a binary classification problem for each pair of genes. A statistical classifier is trained to recognize the relationships between the activation profiles of gene pairs. This approach has been proven to outperform previous unsupervised methods. However, the supervised approach raises open questions. In particular, although known regulatory connections can safely be assumed to be positive training examples, obtaining negative examples is not straightforward, because definite knowledge is typically not available that a given pair of genes do not interact. A recent advance in research on data mining is a method capable of learning a classifier from only positive and unlabeled examples, that does not need labeled negative examples. Applied to the reconstruction of gene regulatory networks, we show that this method significantly outperforms the current state of the art of machine learning methods. We assess the new method using both simulated and experimental data, and obtain major performance improvement. Compared to unsupervised methods for gene network inference, supervised methods are potentially more accurate, but for training they need a complete set of known regulatory connections. A supervised method that can be trained using only positive and unlabeled data, as presented in this paper, is especially beneficial for the task of inferring gene regulatory networks, because only an incomplete set of known regulatory connections is available in public databases such as RegulonDB, TRRD, KEGG, Transfac, and IPA.
DOI: 10.1093/bioinformatics/btl441
发表时间: 2006-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Wang, Chunlin;Ding, Chris;Holbrook, Stephen R.
通讯作者: Holbrook, Stephen R.
DOI: 10.1186/1471-2105-7-s1-s7
发表时间: 2006-03-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Margolin AA;Nemenman I;Basso K;Wiggins C;Stolovitzky G;Dalla Favera R;Califano A
通讯作者: Califano A
DOI: 10.1038/msb4100120
发表时间: 2007
影响因子: 9.9
作者:
通讯作者: --
DOI: 10.1007/s10994-007-5018-6
发表时间: 2007-10-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Lin, Hsuan-Tien;Lin, Chih-Jen;Weng, Ruby C.
通讯作者: Weng, Ruby C.
DOI: 10.1002/9783527622818.ch5
发表时间: 2008-01-01
期刊: ANALYSIS OF MICROARRAY DATA: A NETWORK-BASED APPROACH
影响因子: --
作者:
Grzegorczyk, Marco;Husmeier, Dirk;Werhli, Adriano V.
通讯作者: Werhli, Adriano V.