Multi-documents Automatic Abstracting based on text clustering and semantic analysis

Multi-documents Automatic Abstracting based on text clustering and semantic analysis
复制标题

DOI:
10.1016/j.knosys.2009.06.010
复制
发表时间:
2009-08
期刊:
Knowl. Based Syst.
影响因子:
--
通讯作者:
Qing-lin Guo;Ming Zhang
Qing-lin Guo;Ming Zhang
中科院分区:
其他
文献类型:
--
作者:
Qing-lin Guo;Ming Zhang

文献摘要

相似文献

针对现有的多文档自动文摘方法的不足,提出了一种基于文本聚类和语义分析的多文档自动文摘的实现方法。该方法利用语义分析,实现了多文档的自动文摘。提出了基于标题和段首句的二次自动分词算法.其准确率和召回率均在95%以上。针对塑料领域的一个特定领域,实现了一个自动文摘系统TCAAS。多文档自动文摘的查准率和查全率均在75%以上。实验证明,利用该方法开发领域自动文摘系统是可行的,具有进一步深入研究的价值。
A method of realization of multi-documents Automatic Abstracting based on text clustering and semantic analysis is brought forward, aimed at overcoming shortages of some current methods about multi-documents. The method makes use of semantic analysis and can realize Automatic Abstracting of multi-documents. The algorithm of twice word segmentation based on the title and first-sentences in paragraphs is brought forward. Its precision and recall is above 95%. For a specific domain on plastics, an Automatic Abstracting system named TCAAS is implemented. The precision and recall of multi-document’s Automatic Abstracting is above 75%. And experiments do prove that it is feasible to use the method to develop a domain Automatic Abstracting system, which is valuable for further study in more depth.