Rescoring Confusion Networks for Keyword Search

Rescoring Confusion Networks for Keyword Search
复制标题

重新评分关键字搜索的混淆网络

DOI:
10.1109/icassp.2014.6854975
复制
发表时间:
2014
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Julia Hirschberg
Julia Hirschberg
中科院分区:
--
文献类型:
--
作者:
Víctor Soto;Erica Cooper;L. Mangu;A. Rosenberg;Julia Hirschberg

文献摘要

被引文献

相似文献

本文提出了一种两阶段级联方案,用于在低资源语言环境下对混淆网络(CNS)进行关键词搜索。在第一阶段,我们利用大量的词汇、语音、虚警和结构特征对CN进行重新评分,以提高1-Best假设的错误率。使用一个等级学习支持向量机分类器,我们在粤语、他加路语、土耳其语、普什图语和越南语上获得了0.54%到2.84%的WER收益。在第二阶段,我们从重新评分的CN中生成关键词命中,并使用Logistic回归来检测真实命中和错误警报。我们将这些与未重新评分的CN生成的命中率进行比较,通过使用提到的特征并包括他加路语、土耳其语和普什图语上的声学和韵律特征,在MTWV指标上获得了0.45%到0.9%的收益。
We introduce a two-stage cascaded scheme to rescore Confusion Networks (CNs) for Keyword Search in the context of Low-Resource Languages. In the first stage we rescore the CN to improve the error rate of the 1-best hypothesis using a large number of lexical, phonetic, false alarms and structural features. Using a rank learning Support Vector Machine classifier, we obtain WER gains between 0.54% and 2.84% on Cantonese, Tagalog, Turkish, Pashto and Vietnamese. In the second stage we generate keyword hits from the rescored CN and use logistic regression to detect true hits and false alarms. We compare these to hits generated from the unrescored CN and obtain gains between 0.45% and 0.9% on the MTWV metric by using the mentioned features and including acoustic and prosodic features on Tagalog, Turkish and Pashto.