Using rule-based machine learning for candidate disease gene prioritization and sample classification of cancer gene expression data.

Using rule-based machine learning for candidate disease gene prioritization and sample classification of cancer gene expression data.
复制标题

DOI:
10.1371/journal.pone.0039932
复制
发表时间:
2012
期刊:
影响因子:
3.7
通讯作者:
Krasnogor N
Krasnogor N
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Glaab E;Bacardit J;Garibaldi JM;Krasnogor N

文献摘要

参考文献

被引文献

相似文献

微阵列数据分析已被证明是研究癌症和遗传疾病的有效工具。虽然经典的机器学习技术已经成功地应用于寻找信息基因和预测新样本的类别标签,但微阵列分析的常见限制,如小样本量,大属性空间和高噪声水平仍然限制了其科学和临床应用。提高预测模型的可解释性,同时保持高准确性,将有助于更有效地利用微阵列数据中的信息内容。为了这个目的,我们评估我们的基于规则的进化机器学习系统,BioHEL和GAssist,在三个公共的微阵列癌症数据集,获得简单的基于规则的样本分类模型。基于三种不同特征选择算法的其他基准微阵列样本分类器的比较表明,这些进化学习技术可以与支持向量机等最先进的方法竞争。所获得的模型达到90%以上的准确性,在两级外部交叉验证,与附加值,方便解释,仅使用简单的组合,如果,然后,否则规则。作为另一个好处,文献挖掘分析表明,从BioHEL的分类规则集提取的信息基因的优先级可以优于从传统的集成功能选择获得的基因排名在相关疾病术语和排名靠前的基因的标准化名称之间的逐点互信息。
Microarray data analysis has been shown to provide an effective tool for studying cancer and genetic diseases. Although classical machine learning techniques have successfully been applied to find informative genes and to predict class labels for new samples, common restrictions of microarray analysis such as small sample sizes, a large attribute space and high noise levels still limit its scientific and clinical applications. Increasing the interpretability of prediction models while retaining a high accuracy would help to exploit the information content in microarray data more effectively. For this purpose, we evaluate our rule-based evolutionary machine learning systems, BioHEL and GAssist, on three public microarray cancer datasets, obtaining simple rule-based models for sample classification. A comparison with other benchmark microarray sample classifiers based on three diverse feature selection algorithms suggests that these evolutionary learning techniques can compete with state-of-the-art methods like support vector machines. The obtained models reach accuracies above 90% in two-level external cross-validation, with the added value of facilitating interpretation by using only combinations of simple if-then-else rules. As a further benefit, a literature mining analysis reveals that prioritizations of informative genes extracted from BioHEL’s classification rule sets can outperform gene rankings obtained from a conventional ensemble feature selection in terms of the pointwise mutual information between relevant disease terms and the standardized names of top-ranked genes.
DOI: 10.1158/0008-5472.can-05-4279
发表时间: 2006-08-15
期刊: CANCER RESEARCH
影响因子: 11.2
作者:
Acuff, Heath B.;Sinnamon, Mark;Matrisian, Lynn M.
通讯作者: Matrisian, Lynn M.
DOI: 10.1093/jn/132.11.3451s
发表时间: 2002-11-01
影响因子: 4.2
作者:
Bray, GA
通讯作者: Bray, GA
DOI: 10.1093/bib/bb1016
发表时间: 2007-01-01
影响因子: 9.5
作者:
Boulesteix, Anne-Laure;Strimmer, Korbinian
通讯作者: Strimmer, Korbinian
DOI: 10.1186/bcr1512
发表时间: 2006
期刊: Breast cancer research : BCR
影响因子: --
作者:
Alexe G;Alexe S;Axelrod DE;Bonates TO;Lozina II;Reiss M;Hammer PL
通讯作者: Hammer PL
DOI: 10.1182/blood-2008-07-168096
发表时间: 2009-03-19
期刊: BLOOD
影响因子: 20.3
作者:
Chetaille, Bruno;Bertucci, Francois;Xerri, Luc
通讯作者: Xerri, Luc