Multilingual Controllable Voice Privacy (VoiPy)
Multilingual Controllable Voice Privacy (VoiPy)
批准号:
533241795
负责人:
Professor Dr. Ngoc Thang Vu
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
自动语音处理支持有用的应用程序,如语音助理或自动转录服务。然而,语音信号包含的数据比任务通常所需的数据要多,例如关于说话者的非语言信息。随着以说话人识别或语音情感识别为目标的分类模型的发展,这些信息往往在说话人不知情的情况下构成了严重的隐私威胁。服务提供商可能会滥用数据进行扬声器分析,或者在数据传输或存储过程中落入攻击者手中。因此,在录音后和进一步处理前立即对音频进行匿名化(以保护语音隐私)增加了其相关性和重要性,特别是在欧盟引入通用数据保护条例(GDPR)之后。这个想法是修改语音记录,这样原始说话者和音频之间的联系就被破坏了,例如,通过使用不同说话者的声音。理论上,如果匿名化成功,则不必采取进一步的行动来保护数据中的其他个人属性,如说话者特征(例如,性别、种族出身)或说话者国家(例如,健康状态、情绪)。在这个项目中,我们提出了一个语音隐私框架,让用户控制哪些个人信息,包括但不限于他们的身份,他们希望在与外部服务共享之前在他们的讲话中保护。用户可以灵活地设置针对与其简档相关的不同发言者属性的匿名化请求(例如,年龄、性别)和状态(例如,情绪)。默认情况下,我们不会匿名化所有个人信息,因为根据匿名音频的应用,某些属性需要以未经修改的形式保留。特别是由于智能设备的用户往往不知道他们的隐私受到保护或侵犯,我们专注于为用户提供尽可能多的可控性和透明度,同时保持可用性和有效性。此外,为了将隐私支持扩展到非英语使用者,我们在我们提出的系统中包含了一个多语言切换器,该切换器选择语言相关组件并通知多语言组件有关输入语言。在本提案中,我们将重点关注德国(官方或外国)使用的主要语言,但提出了一种可扩展到其他语言的方法。
英文摘要
Automatic speech processing enables useful applications, like speech assistants or automatic transcription services. However, a speech signal contains more data than is usually needed for the task, such as paralinguistic information about the speaker. As classification models with objectives like speaker identification or speech emotion recognition become advanced, this information poses serious privacy threats, often without the speaker’s knowledge. Service providers might abuse the data for speaker profiling, or it might fall into the hands of attackers during data transmission or storage. Therefore, anonymizing the audio - to preserve voice privacy - immediately after recording and before further processing has increased its relevance and importance, especially after the introduction of the General Data Protection Regulation (GDPR) of the European Union. The idea is to modify speech recordings such that the link between the original speaker and the audio is destroyed, for instance, by using the voice of a different speaker. In theory, if the anonymization has been successful, no further actions have to be taken to protect other personal attributes in the data, like speaker traits (e.g., gender, ethnic origin) or speaker states (e.g., health state, emotions). In this project, we propose a voice privacy framework that gives users control over which personal information, including but not limited to their identity, they want to protect in their speech before sharing it with an external service. The user can flexibly set anonymization requests for different speaker attributes related to their profile (e.g., age, gender) and state (e.g., emotion). We will not anonymize all personal information by default because, depending on the application of the anonymized audio, some attributes need to be preserved in an unmodified form. Especially since users of smart devices often do not know about the protection or violation of their privacy, we focus on providing as much controllability and transparency to the user as possible while keeping usability and effectiveness. Furthermore, to extend privacy support to non-English speakers, we include a multilingual switch in our proposed system that selects language-dependent components and informs multilingual components about the input language. In this proposal, we will focus on the major languages spoken in Germany (official or foreign) but propose a method that is extendable to other languages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金