Textual data mining of service center call records

Textual data mining of service center call records
复制标题

服务中心通话记录文本数据挖掘

DOI:
--
复制
发表时间:
2000
期刊:
Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
R. Goldman
R. Goldman
中科院分区:
--
文献类型:
--
作者:
P. Tan;H. Blau;S. Harp;R. Goldman

文献摘要

被引文献

相似文献

在该项目中,我们开发了一种从包含xedform和自由文本字段的数据库中提取有用信息的技术。数据挖掘的当前技术状态是只处理xed格式数据的技术(模式识别,机器学习的分类算法)和为自由格式文本设计的技术(信息检索)之间的分裂。在这两个研究领域中已经开发了先进的知识检索技术,但是仍然缺乏能够对包含这两种数据的记录进行分类或聚类的系统。具体而言,我们检查了霍尼韦尔服务中心的数据库记录,以提取有关不同类型服务请求的预期成本的信息。我们的目标是测试这样一个假设,即从自由文本电子标签中整合信息将为这些记录提供更好的分类;在这种情况下,更好地预测服务呼叫的成本。在我们的工作中,我们已经集成了特征提取和聚类技术从信息检索与分类算法从机器学习,以分类的混合ELD。我们的初步结果表明,将自由形式的文本可能会导致更好的分类模型。
In this project, w e dev eloped a technique for extracting useful information from databases that contain both xedformat and free-text elds. The present state of the art in data mining is a schism betw een tec hniques that handle only xed-format data (pattern recognition, classi cation algorithms from machine learning), and techniques designed for free-form text (information retrieval). Advanced knowledge disco very technologies ha ve been developed in both research areas, but systems that can categorize or cluster records containing both kinds of data are still lacking. Speci cally, we examined database records from a Honeywell service cen ter to extract information about the expected cost of di erent kinds of service requests. Our goal was to test the h ypothesis that incorporating information from free-text elds would provide a better categorization of these records; in this case, better predictions of the cost of the service call. In our w ork, we have integrated feature extraction and clustering techniques from information retrieval with classi cation algorithms from machine learning in order to categorize the hybrid elds. Our preliminary results suggested that incorporating free-form text could potentially induce better classi cation models.