Combining structural and citation-based evidence for text classification

Combining structural and citation-based evidence for text classification
复制标题

结合结构证据和基于引文的证据进行文本分类

DOI:
--
复制
发表时间:
2004
期刊:
International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Marco Cristo
Marco Cristo
中科院分区:
--
文献类型:
--
作者:
Baoping Zhang;Marcos André Gonçalves;Weiguo Fan;Yuxin Chen;E. Fox;P. Calado;Marco Cristo

文献摘要

被引文献

相似文献

本文讨论了基于引用的信息和结构内容(例如,标题、摘要)可以被组合以改进将文本文档分类为预定义的类别。我们评估不同的措施,从引文结构和结构的集合内容的相似性,并确定如何将它们融合,以提高分类的有效性。为了发现最佳的融合框架,我们采用遗传编程(GP)技术。我们的实证实验使用的文件从ACM数字图书馆和ACM计算分类系统表明,我们可以发现相似性功能,工作比孤立地使用证据,其综合性能通过简单的多数表决是可比的支持向量机分类器。
This paper discusses how citation-based information and structural content (e.g., title, abstract) can be combined to improve classification of text documents into predefined categories. We evaluate different measures of similarity derived from the citation structure and the structural content of the collection, and determine how they can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our empirical experiments using documents from the ACM Digital Library and the ACM Computing Classification System show that we can discover similarity functions that work better than using evidence in isolation and whose combined performance through a simple majority voting is comparable to that of Support Vector Machine classifiers.