An empirical investigation of sparse log-linear models for improved dialogue act classification

An empirical investigation of sparse log-linear models for improved dialogue act classification
复制标题

用于改进对话行为分类的稀疏对数线性模型的实证研究

DOI:
10.1109/icassp.2013.6639287
复制
发表时间:
2013
期刊:
2013 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Alexander I. Rudnicky
Alexander I. Rudnicky
中科院分区:
--
文献类型:
--
作者:
Yun;William Yang Wang;Alexander I. Rudnicky

文献摘要

被引文献

相似文献

以前的对话行为分类工作主要集中在密集的生成和歧视模型。然而,由于自动语音识别(ASR)的输出往往是嘈杂的,密集的模型可能会产生有偏的估计和过拟合的训练数据。在本文中,我们研究稀疏建模方法,以提高对话行为分类,因为稀疏模型保持一个紧凑的特征空间,这是强大的噪声。为了测试这一点,我们研究了各种元素方面的频率收缩模型,如套索,脊和弹性网,以及结构化稀疏模型和层次稀疏模型,嵌入依赖结构和局部特征之间的相互作用。在我们对真实世界数据集的实验中,当使用混淆网络特征增强N个最佳单词和音素级别的ASR假设时,我们最好的稀疏对数线性模型比基于规则的基线获得了19.7%的相对改善,比传统的非稀疏对数线性模型显著改善了3.7%,并且比最先进的SVM模型性能高出2.2%。
Previous work on dialogue act classification have primarily focused on dense generative and discriminative models. However, since the automatic speech recognition (ASR) outputs are often noisy, dense models might generate biased estimates and overfit to the training data. In this paper, we study sparse modeling approaches to improve dialogue act classification, since the sparse models maintain a compact feature space, which is robust to noise. To test this, we investigate various element-wise frequentist shrinkage models such as lasso, ridge, and elastic net, as well as structured sparsity models and a hierarchical sparsity model that embed the dependency structure and interaction among local features. In our experiments on a real-world dataset, when augmenting N-best word and phone level ASR hypotheses with confusion network features, our best sparse log-linear model obtains a relative improvement of 19.7% over a rule-based baseline, a 3.7% significant improvement over a traditional non-sparse log-linear model, and outperforms a state-of-the-art SVM model by 2.2%.