Unsupervised multilingual sentence boundary detection

Unsupervised multilingual sentence boundary detection
复制标题

DOI:
10.1162/coli.2006.32.4.485
复制
发表时间:
2006-12-01
影响因子:
9.3
通讯作者:
Strunk, Jan
Strunk, Jan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kiss, Tibor;Strunk, Jan

文献摘要

被引文献

相似文献

在这篇文章中,我们提出了一种独立于语言的,无监督的方法来检测句子边界。它是基于这样的假设,即一旦缩写被识别,在确定句子边界时的大量歧义就可以消除。而不是依赖于正字法的线索,所提出的系统是能够检测缩写与高精度使用三个标准,只需要有关的候选类型本身的信息,是独立的上下文:缩写可以被定义为一个非常紧密的搭配组成的截断词和最后一个期间,缩写通常是短的,缩写有时包含内部期间。我们还显示了潜在的搭配证据的句子边界消歧的其他两个重要的子任务,即,检测的首字母和序数。所提出的系统已被广泛测试11种不同的语言和不同的文本类型。它取得了良好的效果,没有任何进一步的修改或特定语言的资源。我们评估其性能对三个不同的基线,并将其与文献中提出的句子边界检测的其他系统进行比较。
In this article, we present a language-independent, unsupervised approach to sentence boundary detection. It is based on the assumption that a large number of ambiguities in the determination of sentence boundaries can be eliminated once abbreviations have been identified. Instead of relying on orthographic clues, the proposed system is able to detect abbreviations with high accuracy using three criteria that only require information about the candidate type itself and are independent of context: Abbreviations can be defined as a very tight collocation consisting of a truncated word and a final period, abbreviations are usually short, and abbreviations sometimes contain internal periods. We also show the potential of collocational evidence for two other important subtasks of sentence boundary disambiguation, namely, the detection of initials and ordinal numbers. The proposed system has been tested extensively on eleven different languages and on different text genres. It achieves good results without any further amendments or language-specific resources. We evaluate its performance against three different baselines and compare it to other systems for sentence boundary detection proposed in the literature.