Building a speech recognition system with privacy identification information based on Google Voice for social robots

Building a speech recognition system with privacy identification information based on Google Voice for social robots
复制标题

DOI:
10.1007/s11227-022-04487-3
复制
发表时间:
2022-04
期刊:
The Journal of Supercomputing
影响因子:
--
通讯作者:
Pei-Chun Lin;Benjamin Yankson;Vishal Chauhan;Manabu Tsukada
Pei-Chun Lin;Benjamin Yankson;Vishal Chauhan;Manabu Tsukada
中科院分区:
其他
文献类型:
--
作者:
Pei-Chun Lin;Benjamin Yankson;Vishal Chauhan;Manabu Tsukada

文献摘要

相似文献

目前,市场上出现了许多智能音箱,甚至社交机器人,以帮助人们的生活变得更加方便。通常,人们使用智能扬声器来查看日常日程安排或控制家中的家用电器。许多社交机器人还包括智能扬声器。它们具有用于语音控制机器的共同特性。无论在哪里安装和使用智能扬声器,当人们开始与语音设备对话时,都暴露了安全或隐私风险。因此,本文希望构建一个包含隐私识别信息(PII)的语音识别系统。我们称之为SR-PII系统。我们使用谷歌发布的谷歌人工智能(AIY)语音工具包来构建一个简单,智能的对话扬声器,并包括我们的SR-PII系统。在我们的实验中,我们在三种环境(安静、噪音和播放音乐)中测试了SR的准确性和隐私设置的可靠性。我们还在实验中检查了云响应和说话人响应时间。结果表明,在云环境下,说话人的响应时间约为3.74 s,在云环境下,说话人的响应时间约为9.04 s。我们还展示了说话者的响应准确性,该系统在三种环境中成功地阻止了SR-PII系统的个人信息。在安静的房间中,扬声器的平均响应时间约为8.86秒,平均精度为93%;在嘈杂的环境中,扬声器的平均响应时间约为9.18秒,平均精度为89%;在播放音乐的环境中,扬声器的平均响应时间约为9.62秒,平均精度为90%。我们得出结论,SR-PII系统可以保护隐私信息,而影响说话者响应速度的最重要因素是网络连接状态。我们希望人们可以通过我们的实验,在构建社交机器人和安装SR-PII系统以保护用户个人身份信息方面有一些指导方针。
Currently, many smart speakers, even social robots, appear on the market to help people's lives become more convenient. Usually, people use smart speakers to check their daily schedule or control home appliances in their house. Many social robots also include smart speakers. They have the common property of being used in voice control machines. Regardless of where the smart speaker is installed and used, when people start a conversation with voice equipment, a security or privacy risk is exposed. Hence, we want to build a speech recognition (SR) that contains the privacy identification information (PII) system in this paper. We call this the SR-PII system. We used a Google Artificial-Intelligence-Yourself (AIY) Voice Kit released from Google to build a simple, smart dialog speaker and included our SR-PII system. In our experiments, we test SR accuracy and the reliability of privacy settings in three environments (quiet, noise, and playing music). We also examine the cloud response and speaker response times during our experiments. The results show that the speaker response is approximately 3.74 s in the cloud environment and approximately 9.04 s from the speaker. We also showed the response accuracy of the speaker, which successfully prevented personal information with the SR-PII system in three environments. The speaker has a response mean time of approximately 8.86 s with 93% mean accuracy in a quiet room, approximately 9.18 s with 89% mean accuracy in a noisy environment, and approximately 9.62 s with 90% mean accuracy in an environment that plays music. We conclude that the SR-PII system can secure private information and that the most important factor affecting the response speed of the speaker is the network connection status. We hope that people can, through our experiments, have some guidelines in building social robots and installing the SR-PII system to protect users’ personal identification information.