Deep Learning for Prominence Detection In Children’s Read Speech

Deep Learning for Prominence Detection In Children’s Read Speech
复制标题

DOI:
10.1109/icassp43922.2022.9747780
复制
发表时间:
2021-10
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Kamini Sabu;Mithilesh Vaidya;P. Rao
Kamini Sabu;Mithilesh Vaidya;P. Rao
中科院分区:
其他
文献类型:
--
作者:
Kamini Sabu;Mithilesh Vaidya;P. Rao

文献摘要

被引文献

相似文献

语音中感知显著性的检测已经吸引了从基于知识的语言学和声学特征的设计到从诸如音高和强度轮廓的超切分属性的自动特征学习的方法。相反,我们在这里提出了一个系统,直接操作分段的语音波形学习相关的功能,突出的字检测儿童的口语流利性评估。所选择的CRNN(卷积递归神经网络)框架,结合了词级特征和序列信息,被发现受益于感知激励的SincNet滤波器作为第一个卷积层。我们进一步探讨了在不同的多任务结构下,短语边界和突显的韵律事件之间的语言关联的好处。超越先前报道的性能在同一数据集上的随机森林集成预测训练精心挑选的手工制作的声学特征,我们进一步评估可能的互补信息,从手工制作的声学和预先训练的词汇功能。
The detection of perceived prominence in speech has attracted approaches ranging from the design of knowledge-based linguistic and acoustic features to the automatic feature learning from suprasegmental attributes such as pitch and intensity contours. We present here, in contrast, a system that operates directly on segmented speech waveforms to learn features relevant to prominent word detection for children’s oral fluency assessment. The chosen CRNN (convolutional recurrent neural network) framework, incorporating both word-level features and sequence information, is found to benefit from the perceptually motivated SincNet filters as the first convolutional layer. We further explore the benefits of the linguistic association between the prosodic events of phrase boundary and prominence with different multi-task architectures. Surpassing the previously reported performance on the same dataset of a random forest ensemble predictor trained on carefully chosen hand-crafted acoustic features, we evaluate further the possibly complementary information from hand-crafted acoustic and pre-trained lexical features.