Introducing the consensus modeling concept in genetic algorithms: Application to interpretable discriminant analysis

Introducing the consensus modeling concept in genetic algorithms: Application to interpretable discriminant analysis
复制标题

DOI:
10.1021/ci050529l
复制
发表时间:
2006-09-25
影响因子:
5.6
通讯作者:
Greenidge, Paulette A.
Greenidge, Paulette A.
中科院分区:
化学2区
文献类型:
--
作者:
Ganguly, Milan;Brown, Nathan;Greenidge, Paulette A.

文献摘要

被引文献

相似文献

一种进化统计学习方法被应用于根据药物的生物靶点对药物进行分类,并区分口服和非口服药物的汇编。重点不仅放在模型的预测效果上,而且放在它们的可解释性上。在增强以前的研究中,模型权重的一致性在几个运行的遗传算法被认为是与生产可理解的模型的目标。通过这种方法,描述符和它们的范围,最有助于类歧视被确定。选择能够捕获正在训练的类的平均描述符属性的bin步长,可以提高模型的可解释性和区分能力。这些模型的性能,一致性和鲁棒性进一步增强,通过使用两种新的方法,减少个别解决方案之间的差异:共识和剪接建模。最后,遗传算法区分活性类的能力进行了比较,相似性搜索方法,而朴素贝叶斯分类器和支持向量机应用于区分口服和非口服药物。
An evolutionary statistical learning method was applied to classify drugs according to their biological target and also to discriminate between a compilation of oral and nonoral drugs. The emphasis was placed not only on how well the models predict but also on their interpretability. In an enhancement to previous studies, the consistency of the model weights over several runs of the genetic algorithm was considered with the goal of producing comprehensible models. Via this approach, the descriptors and their ranges that contribute most to class discrimination were identified. Selecting a bin step size that enables the average descriptor properties of the class being trained to be captured improves the interpretability and discriminatory power of a model. The performance, consistency, and robustness of such models were further enhanced by using two novel approaches that reduce the variability between individual solutions: consensus and splice modeling. Finally, the ability of the genetic algorithm to discriminate between activity classes was compared with a similarity searching method, while naive Bayes classifiers and support vector machines were applied in discriminating the oral and nonoral drugs.