Can Prosody Aid the Automatic Processing of Multi-Party Meetings? Evidence from Predicting Punctuation, Disfluencies, and Overlapping Speech

Can Prosody Aid the Automatic Processing of Multi-Party Meetings? Evidence from Predicting Punctuation, Disfluencies, and Overlapping Speech
复制标题

Prosody 可以帮助自动处理多方会议吗?

DOI:
--
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
D. Baron
D. Baron
中科院分区:
--
文献类型:
--
作者:
Elizabeth Shriberg;A. Stolcke;D. Baron

文献摘要

被引文献

相似文献

我们研究韵律的概率建模是否可以帮助处理多方会议所必需的各种自动标记任务。任务1,自动标点符号,旨在对句子边界和不流利进行分类。任务2,跳点,预测背景说话者开始说话的前景语音的位置;任务3,跳入单词,检查他们用来说话的语音特征。数据来自ICSI会议记录器语料库。为了推断内在线索,分析是基于近距离说话的麦克风信号和识别器强制对准。作为词级线索的宽泛基线,我们将韵律模型与给定真实单词的语言模型的韵律模型进行比较。任务1的实验结果表明,与作弊语言模型相比,韵律模型减少了10%的分类错误;此外,当该任务以在线模式运行时,韵律模型的退化程度小于语言模型。对于任务2,语言模型没有提供任何信息,而韵律模型比随机模型减少了13%的熵。对于任务3,韵律模型比随机模型减少了25%的熵。分析还显示了有趣的韵律模式,不同的任务会有所不同。任务1使用与Switchboard(但不是广播新闻)数据类似的提示。任务2预测的跳转点在韵律上看起来像句子边界,但实际上并不是这样的边界。任务3显示,与在沉默状态下开始讲话相比,说话者在开始与他人交谈时会“提高”自己的声音。这些结果证明韵律建模可以用于会议的自动处理。讨论了进一步的结果和对未来自动会议处理系统的影响。
We investigate whether probabilistic modeling of prosody can aid various automatic labeling tasks essential for processing of multi-party meetings. Task 1, automatic punctuation, seeks to classify sentence boundaries and disfluencies. Task 2, jumpin points, predicts locations within foreground speech at which background speakers start talking; Task 3, jump-in words, examines characteristics of the speech they use to do so. Data are from the ICSI Meeting Recorder corpus. To infer inherent cues, analyses are based on close-talking microphone signals and recognizer forced alignments. As a generous baseline for word-level cues, we compare prosodic models to those of a language model given the true words. Results for Task 1 show prosody reduces classification error by 10% relative over the cheating language model; furthermore when this task is run in “online” mode the prosodic model degrades less than does the language model. For Task 2, the language model provides no information, while the prosodic model reduces entropy by 13% over chance. For Task 3, a prosodic model reduces entropy by 25% over chance. Analyses also show interesting prosodic patterns, which differ over tasks. Task 1 uses cues similar to those for Switchboard (but not Broadcast News) data. Task 2 predicts jump-in points that look prosodically like sentence boundaries but that are not actually such boundaries. And Task 3 shows that speakers “raise” their voice when starting during another’s talk, compared to starting during silence. These results provide evidence that prosodic modeling can be of use for the automatic processing of meetings. Further results and implications for future automatic meeting processing systems are discussed.