A Method for Extracting Important Segments from Documents Using Support Vector Machines

A Method for Extracting Important Segments from Documents Using Support Vector Machines
复制标题

一种利用支持向量机从文档中提取重要片段的方法

DOI:
10.1527/tjsai.21.330
复制
发表时间:
2006
影响因子:
--
通讯作者:
A. Utsumi
A. Utsumi
中科院分区:
--
文献类型:
--
作者:
Daisuke Suzuki;A. Utsumi

文献摘要

被引文献

相似文献

本文提出了一种基于抽取的自动文摘方法。该方法包括两个过程:重要片段提取和句子压缩。在重要片段提取过程中,通过支持向量机将文档中的每个片段分类为重要或不重要。然后,句子压缩过程根据句子的依存结构和支持向量机的分类结果来确定句子的语法上合适的部分以用于摘要。为了测试我们方法的性能,我们使用文本摘要挑战(TSC-1)人工准备的摘要语料库进行了评估实验。实验结果表明,我们的方法比纯片段抽取方法和Lead方法取得了更好的性能,特别是对于只有一部分包含在人类摘要中的句子。对实验结果的进一步分析表明,将句子提取和片段提取相结合的混合方法可能会产生更好的摘要。
In this paper we propose an extraction-based method for automatic summarization. The proposed method consists of two processes: important segment extraction and sentence compaction. The process of important segment extraction classifies each segment in a document as important or not by Support Vector Machines (SVMs). The process of sentence compaction then determines grammatically appropriate portions of a sentence for a summary according to its dependency structure and the classification result by SVMs. To test the performance of our method, we conducted an evaluation experiment using the Text Summarization Challenge (TSC-1) corpus of human-prepared summaries. The result was that our method achieved better performance than a segment-extraction-only method and the Lead method, especially for sentences only a part of which was included in human summaries. Further analysis of the experimental results suggests that a hybrid method that integrates sentence extraction with segment extraction may generate better summaries.