Sentence Boundary Detection and the Problem with the U.S.

Sentence Boundary Detection and the Problem with the U.S.
复制标题

DOI:
10.3115/1620853.1620920
复制
发表时间:
2009-05
期刊:
--
影响因子:
--
通讯作者:
D. Gillick
D. Gillick
中科院分区:
其他
文献类型:
--
作者:
D. Gillick

文献摘要

被引文献

相似文献

句子边界检测被广泛使用,但通常使用过时的工具。我们讨论了是什么让它变得困难,哪些特征是相关的,并提供了一个全面的统计系统,现在可以公开使用,它给出了标准新闻语料库中最著名的错误率:在大约27,000个例子中,我们的系统犯了67个错误,其中23个错误涉及“美国”这个词。
Sentence Boundary Detection is widely used but often with outdated tools. We discuss what makes it difficult, which features are relevant, and present a fully statistical system, now publicly available, that gives the best known error rate on a standard news corpus: Of some 27,000 examples, our system makes 67 errors, 23 involving the word "U.S."