Sample Size and Modeling Accuracy with Decision Tree Based Data Mining Tools

Sample Size and Modeling Accuracy with Decision Tree Based Data Mining Tools
复制标题

基于决策树的数据挖掘工具的样本大小和建模准确性

DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
Bern Carey
Bern Carey
中科院分区:
--
文献类型:
--
作者:
James Morgan;R. Dougherty;Allan Hilchie;Bern Carey

文献摘要

被引文献

相似文献

考虑到与非常大的数据集建模相关的成本和基于决策树的模型的过度拟合问题,基于样本的模型是一个有吸引力的替代方案——前提是基于样本的模型具有接近基于所有可用数据的模型的预测精度。本文给出了决策树模型集的结果,这些决策树模型集是跨渐进式样本量集生成的。这些模型分别使用六种著名的商业数据挖掘工具应用于两组实际客户数据。结果表明,随着样本量的增加,模型精度的提高呈递减趋势。当功率曲线拟合到各种样本量的精度估计时,在样本量达到10,000条记录时,在预期终端的0.5%范围内实现了80%以上的时间精度(理论无限样本的精度)。基于这些结果,拟合渐进样本的功率曲线并使用它来建立适当的样本量似乎是一种有希望的机制,可以支持大型数据集的基于样本的建模。
Given the cost associated with modeling very large datasets and over-fitting issues of decision-tree based models, sample based models are an attractive alternative – provided that the sample based models have a predictive accuracy approximating that of models based on all available data. This paper presents results of sets of decision-tree models generated across progressive sets of sample sizes. The models were applied to two sets of actual client data using each of six prominent commercial data mining tools. The results suggest that model accuracy improves at a decreasing rate with increasing sample size. When a power curve was fitted to accuracy estimates across various sample sizes, more than 80 percent of the time accuracy within 0.5 percent of the expected terminal (accuracy of a theoretical infinite sample) was achieved by the time the sample size reached 10,000 records. Based on these results, fitting a power curve to progressive samples and using it to establish an appropriate sample size appears to be a promising mechanism to support sample based modeling for a large dataset.