Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification

Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification
复制标题

DOI:
10.21437/interspeech.2015-92
复制
发表时间:
2015-09
期刊:
--
影响因子:
--
通讯作者:
Sayaka Shiota;F. Villavicencio;J. Yamagishi;Nobutaka Ono;I. Echizen;T. Matsui
Sayaka Shiota;F. Villavicencio;J. Yamagishi;Nobutaka Ono;I. Echizen;T. Matsui
中科院分区:
其他
文献类型:
--
作者:
Sayaka Shiota;F. Villavicencio;J. Yamagishi;Nobutaka Ono;I. Echizen;T. Matsui

文献摘要

被引文献

相似文献

. 摘要本文提出了一种检测欺骗攻击的新对策框架,以降低自动说话人验证(ASV)系统的脆弱性。最近,ASV系统已经达到了与其他生物识别模式相当的性能。然而,针对这些系统的欺骗技术也取得了巨大进展。使用先进的语音合成和语音转换技术的实验显示了不可接受的错误接受率,并探索了几种新的对抗算法来准确检测欺骗材料。然而,目前提出的应对措施是基于自然语音信号和人工语音信号之间的声学差异,预计在不久的将来会逐渐减少。在本文中,我们关注的是语音活性检测,其目的是验证所呈现的语音信号是否来自一个活生生的人。我们使用流行噪音的现象,即当人的呼吸到达麦克风时发生的失真,作为活着的证据。本文提出了流行噪声检测算法,并通过实验研究表明,这些算法可以用于区分实时语音信号和通过语音合成技术产生的人工语音信号。
. Abstract This paper proposes a novel countermeasure framework to detect spoofing attacks to reduce the vulnerability of automatic speaker verification (ASV) systems. Recently, ASV systems have reached equivalent performances equivalent to those of other biometric modalities. However, spoofing techniques against these systems have also progressed drastically. Experimentation using advanced speech synthesis and voice conversion techniques has showed unacceptable false acceptance rates and several new countermeasure algorithms have been explored to detect spoofing materials accurately. However, the counter-measures proposed so far are based on the acoustic differences between natural speech signals and artificial speech signals, expected to become gradually smaller in the near future. In this paper, we focus on voice liveness detection, which aims to validate whether the presented speech signals originated from a live human. We use the phenomenon of pop noise, which is a distortion that happens when human breath reaches a microphone, as liveness evidence. This paper proposes pop noise detection algorithms and shows through an experimental study that they can be used to discriminate live voice signals from artificial ones generated by means of speech synthesis techniques.