"Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification

"Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification
复制标题

“喂?我在跟谁说话?”

DOI:
10.1109/icassp.2019.8682743
复制
发表时间:
2019
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
S. Tubaro
S. Tubaro
中科院分区:
--
文献类型:
--
作者:
Alessandro Lieto;Daniele Moro;Francesco Devoti;Claudia Parera;V. Lipari;Paolo Bestagini;S. Tubaro

文献摘要

被引文献

相似文献

通过深度学习技术增强的自动语音生成算法可实现日益无缝和即时的人机交互。因此,最新一代的电话呼叫机器人听起来比前几代更具有说服力。该技术的应用在隐私问题(例如,在客户服务中)、欺诈行为(例如,社交黑客)和信任侵蚀(例如,生成虚假对话)方面具有强大的社会影响。由于这些原因,识别说话者的本质(人类还是机器人)至关重要。在本文中,我们提出了一种基于卷积神经网络(CNN)的语音分类算法,该算法能够通过分析短音频摘录来自动分类人类和非人类说话者。我们通过利用填充有各种来源的录音的真实人类语音数据库,并使用基于深度学习的最先进的文本到语音生成器(例如 Google WaveNet)自动生成语音来评估所提出的解决方案的有效性。
Automatic speech generation algorithms, enhanced by deep learning techniques, enable an increasingly seamless and immediate machine-to-human interaction. As a result, the latest generation of phone-calling bots sounds more convincingly human than previous generations. The application of this technology has a strong social impact in terms of privacy issues (e.g., in customer-care services), fraudulent actions (e.g., social hacking) and erosion of trust (e.g., generation of fake conversation). For these reasons, it is crucial to identify the nature of a speaker, as either a human or a bot. In this paper, we propose a speech classification algorithm based on Convolutional Neural Networks (CNNs), which enables the automatic classification of human vs non-human speakers from the analysis of short audio excerpts. We evaluate the effectiveness of the proposed solution by exploiting a real human speech database populated with audio recordings from various sources, and automatically generated speeches using state-of-the-art text-to-speech generators based on deep learning (e.g., Google WaveNet).