AUTOMATED LEARNING OF DECISION RULES FOR TEXT CATEGORIZATION

AUTOMATED LEARNING OF DECISION RULES FOR TEXT CATEGORIZATION
复制标题

DOI:
10.1145/183422.183423
复制
发表时间:
1994-07-01
影响因子:
5.6
通讯作者:
WEISS, SM
WEISS, SM
中科院分区:
计算机科学2区
文献类型:
--
作者:
APTE, C;DAMERAU, F;WEISS, SM

文献摘要

被引文献

相似文献

我们描述了在大型文档集合上使用优化的基于规则的归纳方法进行的广泛实验的结果。这些方法的目标是自动发现可用于一般文档分类或自由文本的个性化过滤的分类模式。先前的报告表明,人工设计的基于规则的系统需要多年的开发努力,已成功构建用于“阅读”文档并为其分配主题。我们表明,机器生成的决策规则看起来与人类的表现相当,同时使用相同的基于规则的表示。与其他机器学习技术相比,路透社收集的关键基准测试结果显示,性能大幅提升,从之前报告的 67% 召回率/精度盈亏平衡点提高到 80.5%。在非常高维的特征空间的背景下,研究了几种方法替代方案,包括通用词典与局部词典,以及二进制特征与频率相关特征。
We describe the results of extensive experiments using optimized rule-based induction methods on large document collections. The goal of these methods is to discover automatically classification patterns that can be used for general document categorization or personalized filtering of free text. Previous reports indicate that human-engineered rule-based systems, requiring many man-years of developmental efforts, have been successfully built to ''read'' documents and assign topics to them. We show that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation. In comparison with other machine-learning techniques, results on a key benchmark from the Reuters collection show a large gain in performance, from a previously reported 67% recall/precision breakeven point to 80.5%. In the context of a very high-dimensional feature space, several methodological alternatives are examined, including universal versus local dictionaries, and binary versus frequency-related features.