Topic extraction based on continuous speech recognition in broadcast-news speech

Topic extraction based on continuous speech recognition in broadcast-news speech
复制标题

基于连续语音识别的广播新闻语音主题提取

DOI:
10.1109/asru.1997.659132
复制
发表时间:
1997
期刊:
1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings
影响因子:
--
通讯作者:
S. Furui
S. Furui
中科院分区:
--
文献类型:
--
作者:
K. Ohtsuki;S. Matsunaga;T. Matsuoka;S. Furui

文献摘要

被引文献

相似文献

该论文报道了日本广播新闻演讲中的主题提取。我们研究了使用连续语音识别从广播新闻中提取几个主题词。多个主题词的组合代表了新闻的内容。这比单个词或单个类别更详细、更灵活。主题提取模型显示每个主题词与文章中每个词之间的相关程度。对于文章中的所有单词,从文章中提取总相关得分高的主题词。我们使用五年报纸中的主题词频率和文章中的单词来训练主题提取模型。主题词与文章中的词之间的相关度是根据统计度量(即互信息或 /spl chi//sup 2/ 值)计算的。在识别广播新闻语音的主题提取实验中,我们使用基于 /spl chi//sup 2/ 的模型提取了 5 个主题词,发现其中 75% 与受试者选择的主题词一致。
The paper reports on topic extraction in Japanese broadcast news speech. We studied, using continuous speech recognition, the extraction of several topic words from broadcast news. A combination of multiple topic words represents the content of the news. This is more detailed and more flexible than a single word or a single category. A topic extraction model shows the degree of relevance between each topic word and each word in the articles. For all words in an article, topic words which have high total relevance score are extracted from the article. We trained the topic extraction model with five years of newspapers, using the frequency of topic words taken from headlines and words in articles. The degree of relevance between topic words and words in articles is calculated on the basis of statistical measures, i.e., mutual information or the /spl chi//sup 2/ value. In topic extraction experiments for recognized broadcast news speech, we extracted five topic words using a /spl chi//sup 2/ based model and found that 75% of them agreed with topic words chosen by subjects.