Sample Size and Modeling Accuracy with Decision Tree Based Data Mining Tools
Sample Size and Modeling Accuracy with Decision Tree Based Data Mining Tools
复制标题
基于决策树的数据挖掘工具的样本大小和建模准确性
DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
Bern Carey
中科院分区:
文献类型:
--
作者:
James Morgan;R. Dougherty;Allan Hilchie;Bern Carey
Given the cost associated with modeling very large datasets and over-fitting issues of decision-tree based models, sample based models are an attractive alternative – provided that the sample based models have a predictive accuracy approximating that of models based on all available data. This paper presents results of sets of decision-tree models generated across progressive sets of sample sizes. The models were applied to two sets of actual client data using each of six prominent commercial data mining tools. The results suggest that model accuracy improves at a decreasing rate with increasing sample size. When a power curve was fitted to accuracy estimates across various sample sizes, more than 80 percent of the time accuracy within 0.5 percent of the expected terminal (accuracy of a theoretical infinite sample) was achieved by the time the sample size reached 10,000 records. Based on these results, fitting a power curve to progressive samples and using it to establish an appropriate sample size appears to be a promising mechanism to support sample based modeling for a large dataset.