SOFSEM 2004: Theory and Practice of Computer Science

SOFSEM 2004: Theory and Practice of Computer Science
复制标题

SOFSEM 2004:计算机科学的理论与实践

DOI:
10.1007/b95046
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
Julius Stuller
Julius Stuller
中科院分区:
--
文献类型:
--
作者:
P. E. Boas;J. Pokorný;M. Bieliková;Julius Stuller

文献摘要

被引文献

相似文献

本文描述了在多文档非结构化文本摘要应用中使用熵度量的句子排序技术。该方法是针对特定主题的,并利用简单的语言独立训练框架来计算符号单元的熵。通过将基于熵的分数分配给使用句子相似性的图形表示获得的精简句子集来总结文档集。当应用于相同的数据集时,其性能优于一些常见的统计技术。精度、召回率和 f 分数等常用指标已被修改并用作一组新的指标来比较摘要器的性能。还介绍了这种修改背后的基本原理。实验结果说明了该方法在难以使用特定于语言的词典、翻译器和文档摘要对进行训练的情况下的相关性。
This paper describes a sentence ranking technique using entropy measures, in a multi-document unstructured text summarization application. The method is topic specific and makes use of a simple language independent training framework to calculate entropies of symbol units. The document set is summarized by assigning entropy-based scores to a reduced set of sentences obtained using a graph representation for sentence similarity. The performance is seen to be better than some of the common statistical techniques, when applied on the same data set. Commonly used measures like precision, recall and f-score have been modified and used as a new set of measures for comparing the performance of summarizers. The rationale behind such a modification is also presented. Experimental results are presented to illustrate the relevance of this method in cases where it is difficult to have language specific dictionaries, translators and document-summary pairs for training.