Constructing a Decision Tree for Graph-Structured Data and its Applications

Constructing a Decision Tree for Graph-Structured Data and its Applications
复制标题

DOI:
--
复制
发表时间:
2004-11
期刊:
Fundam. Informaticae
影响因子:
--
通讯作者:
Warodom Geamsakul;Tetsuya Yoshida;K. Ohara;H. Motoda;H. Yokoi;K. Takabayashi
Warodom Geamsakul;Tetsuya Yoshida;K. Ohara;H. Motoda;H. Yokoi;K. Takabayashi
中科院分区:
其他
文献类型:
--
作者:
Warodom Geamsakul;Tetsuya Yoshida;K. Ohara;H. Motoda;H. Yokoi;K. Takabayashi

文献摘要

相似文献

提出了一种基于图的决策树归纳方法(DT-GBI)。在GBI中,通过逐步对扩展(成对分块)在决策树的每个节点处提取子结构(模式),以用作测试的属性。由于属性(特征)是在构造分类器的同时构造的,因此DT-GBI可以被认为是用于特征构造的方法。决策树的预测准确性受到使用哪些属性(模式)以及如何构建它们的影响。在贪婪搜索框架内,采用波束搜索来提取足够好的判别模式。引入了Pessivestry剪枝以避免对训练数据的过拟合。使用DNA数据集进行实验,以观察波束宽度、决策树每个节点处的组块数量和修剪的效果。结果表明,DT-GBI,不使用任何先验领域知识,可以构建一个决策树,是使用领域知识构建的其他分类器相媲美。
Decision tree Graph-Based Induction (DT-GBI) is proposed that constructs a decision tree for graph structured data. Substructures (patterns) are extracted at each node of a decision tree by stepwise pair expansion (pairwise chunking) in GBI to be used as attributes for testing. Since attributes (features) are constructed while a classifier is being constructed, DT-GBI can be conceived as a method for feature construction. The predictive accuracy of a decision tree is affected by which attributes (patterns) are used and how they are constructed. A beam search is employed to extract good enough discriminative patterns within the greedy search framework. Pessimistic pruning is incorporated to avoid overfitting to the training data. Experiments using a DNA dataset were conducted to see the effect of the beam width, the number of chunking at each node of a decision tree, and the pruning. The results indicate that DT-GBI that does not use any prior domain knowledge can construct a decision tree that is comparable to other classifiers constructed using the domain knowledge.