Distributed volunteer computing for solving ensemble learning problems

Distributed volunteer computing for solving ensemble learning problems
复制标题

DOI:
10.1016/j.future.2015.07.010
复制
发表时间:
2016
期刊:
Future Gener. Comput. Syst.
影响因子:
--
通讯作者:
Eugenio Cesario;C. Mastroianni;D. Talia
Eugenio Cesario;C. Mastroianni;D. Talia
中科院分区:
其他
文献类型:
--
作者:
Eugenio Cesario;C. Mastroianni;D. Talia

文献摘要

被引文献

相似文献

志愿者计算范例,沿着定制使用对等通信,最近已经证明能够解决分布式场景中的广泛数据密集型问题。Mining @Homeframework就是基于这些范例的,它已经被实现来运行广泛的分布式数据挖掘应用程序。当整个任务可以被划分为可以并行执行的不同作业时,可以充分利用架构的效率和可扩展性,并且可以重用输入数据,这自然会导致使用数据缓存器。本文探讨了Mining @ Home提供的机会,通过使用装袋方法来处理分类器的发现:使用多个学习器从相同的输入数据中计算模型,以便提取具有高统计精度的最终模型。分析的重点是在一个真实的分布式环境中进行的实验的评估,丰富的仿真评估,以评估非常大的环境,并与分析调查的基础上的等效率的方法。一组广泛的实验允许分析一些异构的情况下,不同的问题大小,这有助于通过适当调整工人的数量和互连域的数量来提高性能。
The volunteer computing paradigm, along with the tailored use of peer-to-peer communication, has recently proven capable of solving a wide area of data-intensive problems in a distributed scenario. TheMining@Homeframework is based on these paradigms and it has been implemented to run a wide range of distributed data mining applications. The efficiency and scalability of the architecture can be fully exploited when the overall task can be partitioned into distinct jobs that may be executed in parallel, and input data can be reused, which naturally leads to the use of data cachers. This paper explores the opportunities offered byMining@Homefor coping with the discovery of classifiers through the use of the bagging approach: multiple learners are used to compute models from the same input data, so as to extract a final model with high statistical accuracy. Analysis focuses on the evaluation of experiments performed in a real distributed environment, enriched with simulation assessment–to evaluate very large environments–and with an analytical investigation based on the iso-efficiency methodology. An extensive set of experiments allowed to analyze a number of heterogeneous scenarios, with different problem sizes, which helps to improve the performance by appropriately tuning the number of workers and the number of interconnected domains.