Statistical models for text segmentation

Statistical models for text segmentation
复制标题

DOI:
10.1023/a:1007506220214
复制
发表时间:
1999-02-01
期刊:
影响因子:
7.5
通讯作者:
Lafferty, J
Lafferty, J
中科院分区:
计算机科学3区
文献类型:
--
作者:
Beeferman, D;Berger, A;Lafferty, J

文献摘要

被引文献

相似文献

本文介绍了一种新的统计方法,自动分割成连贯的段的文本。该方法是基于一种技术,逐步建立一个指数模型,以提取与标记的训练文本中的边界的存在相关的功能。这些模型使用两类特征:主题性特征,其以新颖的方式使用自适应语言模型来检测主题的广泛变化,以及线索词特征,其检测特定词的出现,其可以是特定领域的,倾向于在片段边界附近使用。定量和定性地评估我们的方法表明其在两个非常不同的领域中的有效性,《华尔街日报》的新闻文章和电视广播新闻报道的文字记录。在这些领域的定量结果使用一个新的概率动机的错误度量,它结合了自然和灵活的方式精度和召回。该度量用于对不同特征类型的相对贡献进行定量评估,以及与决策树和先前提出的文本分割算法进行比较。
This paper introduces a new statistical approach to automatically partitioning text into coherent segments. The approach is based on a technique that incrementally builds an exponential model to extract features that are correlated with the presence of boundaries in labeled training text. The models use two classes of features: topicality features that use adaptive language models in a novel way to detect broad changes of topic, and cue-word features that detect occurrences of specific words, which may he domain-specific, that tend to be used near segment boundaries, Assessment of our approach on quantitative and qualitative grounds demonstrates its effectiveness in two very different domains, Wall Street Journal news articles and television broadcast news story transcripts. Quantitative results on these domains are presented using a new probabilistically motivated error metric, which combines precision and recall in a natural and flexible way. This metric is used to make a quantitative assessment of the relative contributions of the different feature types, as well as a comparison with decision trees and previously proposed text segmentation algorithms.