Investigating Statistical Techniques for Sentence-Level Event Classification

Investigating Statistical Techniques for Sentence-Level Event Classification
复制标题

DOI:
10.3115/1599081.1599159
复制
发表时间:
2008-08
期刊:
--
影响因子:
--
通讯作者:
Martina Naughton;N. Stokes;J. Carthy
Martina Naughton;N. Stokes;J. Carthy
中科院分区:
其他
文献类型:
--
作者:
Martina Naughton;N. Stokes;J. Carthy

文献摘要

被引文献

相似文献

正确地对描述事件的句子进行分类是许多自然语言应用程序(如问答和摘要)的重要任务。在本文中,我们将事件检测看作一个句子级别的文本分类问题。我们比较了两种方法的性能:支持向量机(SVM)分类器和语言建模(LM)方法。我们还研究了一种基于规则的方法,该方法使用从WordNet派生的手工制作的术语列表。这些术语与给定的事件类型紧密关联,可用于标识描述该类型实例的句子。我们在实验中使用了两个数据集,并在六种不同的事件类型上对每种技术进行了评估。我们的结果表明,在这项任务中,支持向量机的性能一直优于最大似然技术。更有趣的是,我们发现基于人工规则的分类系统是一个非常强大的基线,在六种事件类型中的三种上优于支持向量机。
The ability to correctly classify sentences that describe events is an important task for many natural language applications such as Question Answering (QA) and Summarisation. In this paper, we treat event detection as a sentence level text classification problem. We compare the performance of two approaches to this task: a Support Vector Machine (SVM) classifier and a Language Modeling (LM) approach. We also investigate a rule based method that uses hand crafted lists of terms derived from WordNet. These terms are strongly associated with a given event type, and can be used to identify sentences describing instances of that type. We use two datasets in our experiments, and evaluate each technique on six distinct event types. Our results indicate that the SVM consistently outperform the LM technique for this task. More interestingly, we discover that the manual rule based classification system is a very powerful baseline that outperforms the SVM on three of the six event types.