Effects of End-to-end ASR and Score Fusion Model Learning for Improved Query-by-example Spoken Term Detection
Effects of End-to-end ASR and Score Fusion Model Learning for Improved Query-by-example Spoken Term Detection
复制标题
DOI:
--
复制
发表时间:
2020-12
期刊:
影响因子:
--
通讯作者:
Takumi Kurokawa;A. Kai;Hiroki Kondo
中科院分区:
文献类型:
--
作者:
Takumi Kurokawa;A. Kai;Hiroki Kondo
Query-by-example spoken term detection (STD) systems can make effective use of automatic speech recognition (ASR), especially in situations where the recognition accuracy is high. However, out-of-vocabulary (OOV) problem at the ASR stage has a significant impact on the performance of STD for speech retrieval and can often occur for query terms. Recent studies have shown that end-to-end (E2E) ASR systems can achieve competitive performance compared to conventional DNNHMM-based ASR systems and reduce the impact of OOV problem by adopting output units of characters or subwords. This paper proposes to apply E2E ASR system in an STD method that considers acoustic similarity at sub-phone level, and to combine it with the DNN-HMM-based ASR and auxiliary information by a score fusion method. Experimental results on the NTCIR-12 SpokenQuery& Doc-2 task showed that the STD method using the hybrid CTC/Transformer E2E ASR improved the search performance over the STD method using the DNNHMM-based ASR. The best detection performance was obtained using a score fusion model, demonstrating that combining E2E ASR and auxiliary information with DNN-HMM-based ASR is effective for both known and OOV word queries.