Maximum entropy models with inequality constraints: A case study on text categorization

Maximum entropy models with inequality constraints: A case study on text categorization
复制标题

DOI:
10.1007/s10994-005-0911-3
复制
发表时间:
2005-09-01
期刊:
影响因子:
7.5
通讯作者:
Tsujii, J
Tsujii, J
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kazama, J;Tsujii, J

文献摘要

被引文献

相似文献

数据稀疏或过度拟合是采用机器学习方法的自然语言处理中的一个严重问题。即使对于最大熵(ME)方法来说也是如此,其灵活的建模能力在许多 NLP 任务中比其他概率模型更成功地缓解了数据稀疏性。虽然我们通常使用 ME 方法来估计模型,使其完全满足特征期望的等式约束,但完全满足会导致不良的过拟合,特别是对于稀疏特征,因为从有限数量的训练数据导出的约束总是不确定的。为了控制 ME 估计中的过度拟合,我们建议使用盒型不等式约束,其中等式可能会被违反到反映这种不确定性的某些预定义水平。导出的模型(即不等式 ME 模型)实际上具有对有界参数进行 L-1 范数惩罚的正则化估计。最重要的是,这种正则化估计使模型参数变得稀疏。这可以被认为是自动特征选择,有望进一步提高泛化性能。我们在文本分类数据集上评估了不等式 ME 模型,并证明了它们相对于标准 ME 估计、类似的 ME 模型高斯 MAP 估计以及支持向量机 (SVM) 的优势,后者是最先进的文本分类方法之一。
Data sparseness or overfitting is a serious problem in natural language processing employing machine learning methods. This is still true even for the maximum entropy (ME) method, whose flexible modeling capability has alleviated data sparseness more successfully than the other probabilistic models in many NLP tasks. Although we usually estimate the model so that it completely satisfies the equality constraints on feature expectations with the ME method, complete satisfaction leads to undesirable overfitting, especially for sparse features, since the constraints derived from a limited amount of training data are always uncertain. To control overfitting in ME estimation, we propose the use of box-type inequality constraints, where equality can be violated up to certain predefined levels that reflect this uncertainty. The derived models, inequality ME models, in effect have regularized estimation with L-1 norm penalties of bounded parameters. Most importantly, this regularized estimation enables the model parameters to become sparse. This can be thought of as automatic feature selection, which is expected to improve generalization performance further. We evaluate the inequality ME models on text categorization datasets, and demonstrate their advantages over standard ME estimation, similarly motivated Gaussian MAP estimation of ME models, and support vector machines (SVMs), which are one of the state-of-the-art methods for text categorization.