Identification of age-group from children's speech by computers and humans

Identification of age-group from children's speech by computers and humans
复制标题

计算机和人类从儿童语音中识别年龄组

DOI:
10.21437/interspeech.2014-61
复制
发表时间:
2014
期刊:
Comput. Speech Lang.
影响因子:
--
通讯作者:
P. Jančovič
P. Jančovič
中科院分区:
--
文献类型:
--
作者:
Saeid Safavi;M. Russell;P. Jančovič

文献摘要

被引文献

相似文献

本文介绍了使用 OGI Kids 语料库以及 GMM-UBM、GMM-SVM 和 i-vector 系统对儿童言语进行年龄组识别 (Age-ID) 的结果。通过在 21 个子频带上进行年龄 ID 实验,可以识别包含儿童重要年龄信息的频谱区域。结果表明,5.5 kHz 以上的频率对于 Age-ID 最没有用处。探讨了使用性别无关和性别依赖年龄组模型的效果。 GMM-UBM 和 i-vector 系统的性能明显优于 GMM-SVM 系统。 85.77% 的最佳 Age-ID 性能是通过应用于 5.5 kHz 频带限制语音的 i-vector 系统获得的。还对人类 Age-ID 进行了实验,结果表明人类达不到机器的性能。
This paper presents results on age-group identification (Age-ID) for children’s speech, using the OGI Kids corpus and GMM-UBM, GMM-SVM and i-vector systems. Regions of the spectrum containing important age information for children are identified by conducting Age-ID experiments over 21 frequency sub-bands. Results show that the frequencies above 5.5 kHz are least useful for Age-ID. The effect of using gender-independent and gender-dependent age-group modelling is explored. The GMM-UBM and i-vector systems considerably outperform the GMM-SVM system. The best Age-ID performance of 85.77% is obtained by the i-vector system applied to band-limited speech to 5.5 kHz. Experiments on human Age-ID were also conducted and the results show that the humans do not achieve the performance of the machine.