Improving discriminative sequential learning with rare--but--important associations

Improving discriminative sequential learning with rare--but--important associations
复制标题

DOI:
10.1145/1081870.1081906
复制
发表时间:
2005-08
期刊:
--
影响因子:
--
通讯作者:
X. Phan;Minh Le Nguyen;T. Ho;S. Horiguchi
X. Phan;Minh Le Nguyen;T. Ho;S. Horiguchi
中科院分区:
其他
文献类型:
--
作者:
X. Phan;Minh Le Nguyen;T. Ho;S. Horiguchi

文献摘要

被引文献

相似文献

像条件随机场(CRFs)这样的判别序列学习模型在自然语言处理或信息提取等领域取得了巨大的成功。它们的主要优势是能够捕获各种非独立和重叠的输入特征。然而,一些意想不到的陷阱对模型的性能产生了负面影响;这些陷阱主要来自类别/标签之间的不平衡,不规则现象以及训练数据中的潜在模糊性。本文提出了一种数据驱动的方法,可以通过发现和强调隐藏在训练数据中的罕见但重要的统计关联来处理这种难以预测的数据实例。挖掘的关联然后被纳入这些模型中,以处理困难的例子。在英语短语组块和命名实体识别中的实验结果表明,该方法在准确率上有显著提高。除了技术角度,我们的方法还强调了关联挖掘和统计学习之间的潜在联系,提供了一种替代策略,通过从大型数据集中发现有趣和有用的模式来提高学习性能。
Discriminative sequential learning models like Conditional Random Fields (CRFs) have achieved significant success in several areas such as natural language processing or information extraction. Their key advantage is the ability to capture various non--independent and overlapping features of inputs. However, several unexpected pitfalls have a negative influence on the model's performance; these mainly come from an imbalance among classes/labels, irregular phenomena, and potential ambiguity in the training data. This paper presents a data--driven approach that can deal with such hard--to--predict data instances by discovering and emphasizing rare--but--important associations of statistics hidden in the training data. Mined associations are then incorporated into these models to deal with difficult examples. Experimental results of English phrase chunking and named entity recognition using CRFs show a significant improvement in accuracy. In addition to the technical perspective, our approach also highlights a potential connection between association mining and statistical learning by offering an alternative strategy to enhance learning performance with interesting and useful patterns discovered from large dataset.