Towards automation of knowledge understanding: An approach for probabilistic generative classifiers

Towards automation of knowledge understanding: An approach for probabilistic generative classifiers
复制标题

DOI:
10.1016/j.ins.2016.08.016
复制
发表时间:
2016-11-20
影响因子:
8.1
通讯作者:
Ovaska, Seppo J.
Ovaska, Seppo J.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fisch, Dominik;Gruhl, Christian;Ovaska, Seppo J.

文献摘要

被引文献

相似文献

在数据选择、预处理、转换和特征提取之后,知识提取不是数据挖掘过程的最后一步。因此,有必要了解这些知识,以便有效和有效地应用它。到目前为止,还缺乏支持这一重要步骤的适当技术。这部分是由于对知识的评估往往是高度主观的,例如,关于新奇或实用性等方面。这些方面取决于数据挖掘器的特定知识和要求。然而,有一些方面是客观的,可以采取适当的措施。在这篇文章中,我们专注于分类问题,并使用基于混合密度模型的概率生成分类器,这在数据挖掘应用中非常常见。我们定义了客观的措施来评估这些分类器中包含的规则的信息量,独特性,重要性,歧视性,代表性,不确定性和可验证性。这些措施不仅支持数据挖掘器评估基于这种分类器的数据挖掘过程的结果。正如我们将在示例性案例研究中看到的,它们也可以用于改进数据挖掘过程本身或支持提取的知识的后续应用。(C)2016 Elsevier Inc. All rights reserved.
After data selection, pre-processing, transformation, and feature extraction, knowledge extraction is not the final step in a data mining process. It is then necessary to understand this knowledge in order to apply it efficiently and effectively. Up to now, there is a lack of appropriate techniques that support this significant step. This is partly due to the fact that the assessment of knowledge is often highly subjective, e.g., regarding aspects such as novelty or usefulness. These aspects depend on the specific knowledge and requirements of the data miner. There are, however, a number of aspects that are objective and for which it is possible to provide appropriate measures. In this article we focus on classification problems and use probabilistic generative classifiers based on mixture density models that are quite common in data mining applications. We define objective measures to assess the informativeness, uniqueness, importance, discrimination, representativity, uncertainty, and distinguishability of rules contained in these classifiers numerically. These measures not only support a data miner in evaluating results of a data mining process based on such classifiers. As we will see in illustrative case studies, they may also be used to improve the data mining process itself or to support the later application of the extracted knowledge. (C) 2016 Elsevier Inc. All rights reserved.