Automatic indexing of lecture speech by extracting topic-independent discourse markers

Automatic indexing of lecture speech by extracting topic-independent discourse markers
复制标题

通过提取与主题无关的话语标记来自动索引讲座演讲

DOI:
10.1109/icassp.2002.5743639
复制
发表时间:
2002
期刊:
2002 IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Masahiro Hasegawa
Masahiro Hasegawa
中科院分区:
--
文献类型:
--
作者:
Tatsuya Kawahara;Masahiro Hasegawa

文献摘要

被引文献

相似文献

研究了演讲中分段(子话题)边界的自动检测问题。该方法利用了被定义为语篇制造者的段落的起始话语的特征表达,以及停顿和语言模型信息。基于信息检索技术中的词统计,以完全无监督的方式提取话语标记。统计数据用于选择其他信息所拾取的候选人。实验结果表明,与仅使用暂停信息的简单基线方法相比,该方法具有更好的索引性能(在高查全率下精度更高)。此外,它对语音识别错误具有鲁棒性。
Automatic detection of section (sub-topic) boundaries in lecture speech is addressed. The method makes use of the characteristic expressions used in initial utterances of sections defined as discourse makers, as well as pause and language model information. The discourse markers are derived in a totally unsupervised manner based on word statistics used in the information retrieval technique. The statistics is used to select candidates picked up by other information. Experimental results show that the proposed method realizes better indexing performance (better precision at high recall rates) than the simple baseline method using pause information only. Moreover, it is shown to be robust against speech recognition errors.