Topic segmentation of news speech using word similarity

Topic segmentation of news speech using word similarity
复制标题

利用词相似度进行新闻语音主题分割

DOI:
10.1145/354384.376354
复制
发表时间:
2000
期刊:
Proceedings of the eighth ACM international conference on Multimedia
影响因子:
--
通讯作者:
Y. Ariki
Y. Ariki
中科院分区:
--
文献类型:
--
作者:
S. Takao;J. Ogata;Y. Ariki

文献摘要

被引文献

相似文献

传统的主题分割采用余弦度量作为连续段落之间的相似度。然而,余弦度量存在一个问题,即除非文章中包含完全相同的单词,否则它不能反映相似度。针对这一问题,本文提出了一种通过收集相同主题段的方法,直接从输入数据中自动获取不同词之间的词相似度。在此基础上,提出了一种基于词相似度的文本相似度计算方法。最后,提出了一种无监督模式下基于段落相似度的主题分割方法。
Conventional topic segmentation utilizes cosine measure as the similarity between consecutive passages. However, the cosine measure has a problem that it can not reflect the similarity unless exactly the same words are included in the passages. To solve this problem, in this paper, we propose a method to acquire the word similarity between different words from the input data directly and automatically by managing to collect the same topic sections. Further more, we propose a method to compute the passage similarity based on the word similarity. Finally we propose a method of topic segmentation based on the passage similarity in an unsupervised mode.