Flash: A GP-GPU Ensemble Learning System for Handling Large Datasets

Flash: A GP-GPU Ensemble Learning System for Handling Large Datasets
复制标题

Flash:用于处理大型数据集的 GP-GPU 集成学习系统

DOI:
10.1007/978-3-662-44303-3_2
复制
发表时间:
2014
期刊:
Int. J. High Perform. Syst. Archit.
影响因子:
--
通讯作者:
Una
Una
中科院分区:
--
文献类型:
--
作者:
Ignacio Arnaldo;K. Veeramachaneni;Una

文献摘要

被引文献

相似文献

Flash系统在共享内存桌面上运行基于集成的遗传规划GP符号回归。为了显著减少符号回归所需的大量模型预测的高时间成本,其适应度评估被分配给桌面的GPU。连续GP“实例”在不同的数据子集和随机选择的目标函数上运行。在固定的代数之后收集最佳模型,然后与自适应的输出空间方法融合。一旦学习完成,新的实例启动就会停止。我们证明了Flash的集成策略不仅使GP更加健壮,而且还提供了一种停止学习过程的知情在线手段。Flash使GP能够从由370K个范例和90个特征组成的数据集中学习,在短短50秒内进化出100代以上的1000个个体。
The Flash system runs ensemble-based Genetic Programming GP symbolic regression on a shared memory desktop. To significantly reduce the high time cost of the extensive model predictions required by symbolic regression, its fitness evaluations are tasked to the desktop's GPU. Successive GP "instances" are run on different data subsets and randomly chosen objective functions. Best models are collected after a fixed number of generations and then fused with an adaptive, output-space method. New instance launches are halted once learning is complete. We demonstrate that Flash's ensemble strategy not only makes GP more robust, but it also provides an informed online means of halting the learning process. Flash enables GP to learn from a dataset composed of 370K exemplars and 90 features, evolving a population of 1000 individuals over 100 generations in as few as 50 seconds.