ExSTraCS 2.0: Description and Evaluation of a Scalable Learning Classifier System.

ExSTraCS 2.0: Description and Evaluation of a Scalable Learning Classifier System.
复制标题

DOI:
10.1007/s12065-015-0128-8
复制
发表时间:
2015-09
影响因子:
2.6
通讯作者:
Moore JH
Moore JH
中科院分区:
其他
文献类型:
--
作者:
Urbanowicz RJ;Moore JH

文献摘要

被引文献

相似文献

在这个“大数据”时代,算法的可扩展性是任何机器学习策略的主要关注点。大量潜在的预测属性是生物信息学、遗传流行病学和许多其他领域问题的象征。以前,ExS-TraCS是作为扩展的密歇根式监督学习分类器系统引入的,该系统结合了一组强大的启发式方法,成功地解决了复杂、嘈杂和异构问题领域中的分类、预测和知识发现挑战。虽然密歇根式学习分类器系统是强大而灵活的学习器,但它们并不具有特别的可扩展性。本文首次对ExS-TraCS算法进行了完整的描述,并引入了一种有效的策略来显著提高学习分类器系统的可扩展性。ExSTraCS 2.0解决了以下方面的可扩展性:(1)规则特异性限制,(2)专家知识指导覆盖和突变机制的新方法,以及(3)实现和利用TuRF算法来提高大型数据集中专家知识发现的质量。在复杂的模拟遗传数据集上的性能表明,这些新机制极大地提高了具有20个属性的数据集上的几乎所有性能指标,并使ExSTraCS能够可靠地扩展到相关的200和2000个属性的数据集上。ExSTraCS 2.0还能够可靠地解决6,11,20,37,70和135个多路复用器问题,并且与以前报道的学习迭代相似或更少,使用更小的有限训练集,并且不使用从更简单的多路复用器问题中发现的构建块。此外,通过消除先前的关键运行参数,ExS-TraCS的可用性变得更加简单。
Algorithmic scalability is a major concern for any machine learning strategy in this age of ‘big data’. A large number of potentially predictive attributes is emblematic of problems in bioinformatics, genetic epidemiology, and many other fields. Previously, ExS-TraCS was introduced as an extended Michigan-style supervised learning classifier system that combined a set of powerful heuristics to successfully tackle the challenges of classification, prediction, and knowledge discovery in complex, noisy, and heterogeneous problem domains. While Michigan-style learning classifier systems are powerful and flexible learners, they are not considered to be particularly scalable. For the first time, this paper presents a complete description of the ExS-TraCS algorithm and introduces an effective strategy to dramatically improve learning classifier system scalability. ExSTraCS 2.0 addresses scalability with (1) a rule specificity limit, (2) new approaches to expert knowledge guided covering and mutation mechanisms, and (3) the implementation and utilization of the TuRF algorithm for improving the quality of expert knowledge discovery in larger datasets. Performance over a complex spectrum of simulated genetic datasets demonstrated that these new mechanisms dramatically improve nearly every performance metric on datasets with 20 attributes and made it possible for ExSTraCS to reliably scale up to perform on related 200 and 2000-attribute datasets. ExSTraCS 2.0 was also able to reliably solve the 6, 11, 20, 37, 70, and 135 multiplexer problems, and did so in similar or fewer learning iterations than previously reported, with smaller finite training sets, and without using building blocks discovered from simpler multiplexer problems. Furthermore, ExS-TraCS usability was made simpler through the elimination of previously critical run parameters.