A MFoM learning approach to robust multiclass multi-label text categorization

A MFoM learning approach to robust multiclass multi-label text categorization
复制标题

DOI:
10.1145/1015330.1015361
复制
发表时间:
2004-07
期刊:
Proceedings of the twenty-first international conference on Machine learning
影响因子:
--
通讯作者:
Sheng Gao;Wen Wu;Chin-Hui Lee;Tat-Seng Chua
Sheng Gao;Wen Wu;Chin-Hui Lee;Tat-Seng Chua
中科院分区:
其他
文献类型:
--
作者:
Sheng Gao;Wen Wu;Chin-Hui Lee;Tat-Seng Chua

文献摘要

被引文献

相似文献

本文提出了一种多类文本分类方法。为了充分利用正样本和负样本的优势,引入了一种最大品质因数(MFoM)学习算法来训练高性能的MC分类器。与传统的二进制分类相比,建议的MC计划为每个给定的测试样本的每个类别分配一个统一的评分函数,因此经典的贝叶斯决策规则现在可以应用。由于所有的MC MFoM分类器都是同时训练的,我们希望它们比二进制MFoM分类器更健壮,工作得更好,二进制MFoM分类器是单独训练的,并且已知可以提供最佳的TC性能。路透社-21578 TC任务的实验结果表明,MC MFoM分类器实现了微平均F1值为0.377,这是显着优于0.138,获得与二进制MFoM分类器,为小于4个训练样本的类别。此外,对于所有90个类别,大多数具有较大的训练大小,MC MFoM分类器给出了0.888的微平均F1值,优于使用二进制MFoM分类器获得的0.884。
We propose a multiclass (MC) classification approach to text categorization (TC). To fully take advantage of both positive and negative training examples, a maximal figure-of-merit (MFoM) learning algorithm is introduced to train high performance MC classifiers. In contrast to conventional binary classification, the proposed MC scheme assigns a uniform score function to each category for each given test sample, and thus the classical Bayes decision rules can now be applied. Since all the MC MFoM classifiers are simultaneously trained, we expect them to be more robust and work better than the binary MFoM classifiers, which are trained separately and are known to give the best TC performance. Experimental results on the Reuters-21578 TC task indicate that the MC MFoM classifiers achieve a micro-averaging F1 value of 0.377, which is significantly better than 0.138, obtained with the binary MFoM classifiers, for the categories with less than 4 training samples. Furthermore, for all 90 categories, most with large training sizes, the MC MFoM classifiers give a micro-averaging F1 value of 0.888, better than 0.884, obtained with the binary MFoM classifiers.