A study in machine learning from imbalanced data for sentence boundary detection in speech

A study in machine learning from imbalanced data for sentence boundary detection in speech
复制标题

DOI:
10.1016/j.csl.2005.06.002
复制
发表时间:
2006-10-01
影响因子:
4.3
通讯作者:
Stolcke, Andreas
Stolcke, Andreas
中科院分区:
计算机科学3区
文献类型:
--
作者:
Liu, Yang;Chawla, Nitesh V.;Stolcke, Andreas

文献摘要

被引文献

相似文献

使用句子边界丰富语音识别输出可以提高其人类可读性,并允许下游语言处理模块进行进一步处理。我们构建了一个隐马尔可夫模型(HMM)系统来检测使用韵律和文本信息的句子边界。由于数据中的非句子边界多于句子边界,因此必须构建作为决策树分类器实现的韵律模型,以有效地从不平衡的数据分布中学习。为了解决这个问题,我们研究了各种采样方法和装袋方案。我们进行了一项试点研究,以选择适用于跨两个语料库(会话电话语音和广播新闻语音)的完整 NIST 句子边界评估任务的方法,使用人类转录和识别输出。在试点研究中,当分类错误率作为性能指标时,使用原始训练集在采样方法中实现了最佳性能,而来自不同下采样训练集的多个分类器的集成实现了稍差的性能,但有可能减少计算量。然而,当使用受试者工作特征 (ROC) 或曲线下面积 (AUC) 来衡量性能时,采样方法的性能优于原始训练集。如果句子边界检测输出,这一观察很重要。由下游语言处理模块使用。我们发现装袋可以显着提高每种采样方法的系统性能。当韵律模型与语言模型结合时,这些方法的收益可能会减少,而语言模型是句子检测任务的强大知识源。这。试点研究中发现的模式在完整的 NIST 评估任务中得到了复制。结论可能取决于任务、分类器和知识组合方法。 (c) 2005 Elsevier Ltd. 保留所有权利。
Enriching speech recognition output with sentence boundaries improves its human readability and enables further processing by downstream language processing modules. We have constructed a hidden Markov model (HMM) system to detect sentence boundaries that uses both prosodic and textual information. Since there are more nonsentence boundaries than sentence boundaries in the data, the prosody model, which is implemented as a decision tree classifier, must be constructed to effectively learn from the imbalanced data distribution. To address this problem, we investigate a variety of sampling approaches and a bagging scheme. A pilot study was carried out to select methods to apply to the full NIST sentence boundary evaluation task across two corpora (conversational telephone speech and broadcast news speech), using both human transcriptions and recognition output. In the pilot study, when classification error rate is the performance measure, using the original training set achieves the best performance among the sampling methods, and an ensemble of multiple classifiers from different downsampled training sets achieves slightly poorer performance, but has the potential to reduce computational effort. However, when performance is measured using receiver operating characteristics (ROC) or area under the curve (AUC), then the sampling approaches outperform the original training set. This observation is important if the sentence boundary detection output. is used by downstream language processing modules. Bagging was found to significantly improve system performance for each of the sampling methods. The gain from these methods may be diminished when the prosody model is combined with the language model, which is,a strong knowledge source for the sentence detection task. The. patterns found in the pilot study were replicated in the full NIST evaluation task. The conclusions may be dependent on the task, the classifiers, and the knowledge combination approach. (c) 2005 Elsevier Ltd. All rights reserved.