An Approach to Automation Selection of Decision Tree based on Training Data Set

An Approach to Automation Selection of Decision Tree based on Training Data Set
复制标题

基于训练数据集的决策树自动化选择方法

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
M.Devi
M.Devi
中科院分区:
--
文献类型:
--
作者:
D.Saravana Kumar;N.Ananthi;M.Devi

文献摘要

被引文献

相似文献

在挖掘应用中,具有数百万条记录的非常大的训练数据集是很常见的。对于分类和预测问题,决策树都是非常强大和优秀的技术。为了开发和处理大大小小的训练数据,已经提出了许多决策树构造算法。一些相关算法最适用于大数据集,而另一些算法则适用于小数据集。每种算法都适用于其自己的标准。决策树算法对分类属性和连续属性进行了很好的分类,但它只有效地处理了较小的数据集。对于大型数据集,它会消耗更多的时间。Quest中的监督学习(SLIQ)和决策树的可伸缩并行归纳(Sprint)处理非常大的数据集。但SLIQ要求类标签应该事先在主内存中可用。Sprint最适合大型数据集,它取消了所有这些内存限制。研究了基于训练数据集大小的决策树算法的自动选择问题。该系统首先使用数学度量来准备训练数据集大小。将用可用存储空间检查结果训练集大小问题。如果内存非常充足,则树构建将继续。在对数据进行分类后,对分类器数据集的精度进行估计。该方法的主要优点是系统运行时间较短,避免了内存问题。
mining applications, very large training data sets with several million records are common. Decision trees are very much powerful and excellent technique for both classification and prediction problems. Many decision tree construction algorithms have been proposed to develop and handle large or small training data. Some related algorithms are best for large data sets and some for small data sets. Each algorithm works best for its own criteria. The decision tree algorithms classify categorical and continuous attributes very well but it handles efficiently only a smaller data set. It consumes more time for large datasets. Supervised Learning In Quest (SLIQ) and Scalable Parallelizable Induction of Decision Tree (SPRINT) handles very large datasets. But SLIQ requires that the class labels should be available in main memory beforehand. SPRINT is best suited for large data sets and it removes all these memory restrictions. The research work deals with the automatic selection of decision tree algorithm based on training dataset size. This proposed system first prepares the training dataset size using the mathematical measure. The result training set size problem will be checked with the available memory space. If memory is very sufficient then the tree construction will continue. After the classifying the data, the accuracy of the classifier data set is estimated. The main advantages of the proposed method are that the system takes less time and avoids memory problem.