课题基金 / 基金详情

Multilingual Controllable Voice Privacy (VoiPy)

Multilingual Controllable Voice Privacy (VoiPy)
多语言可控语音隐私 (VoiPy)
批准号:
533241795
负责人:
Professor Dr. Ngoc Thang Vu
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
自动语音处理支持有用的应用程序,如语音助手或自动转录服务。然而,语音信号包含的数据比任务通常需要的数据更多,例如关于说话人的副语言信息。随着以说话人识别或语音情感识别为目标的分类模型的发展,这些信息构成了严重的隐私威胁,通常是在说话人不知情的情况下。服务提供商可能会将数据滥用于说话人分析,或者在数据传输或存储期间落入攻击者手中。因此,在录音之后和进一步处理之前立即匿名音频--以保护语音隐私--增加了其相关性和重要性,特别是在欧洲联盟引入一般数据保护条例(GDPR)之后。其想法是修改语音记录,以便例如通过使用不同说话者的声音来破坏原始说话者和音频之间的链接。理论上,如果匿名化已经成功,就不必采取进一步的行动来保护数据中的其他个人属性,如说话人的特征(例如,性别、民族血统)或说话人的状态(例如,健康状况、情绪)。在这个项目中,我们提出了一个语音隐私框架,该框架允许用户在与外部服务共享之前,控制他们希望在语音中保护哪些个人信息,包括但不限于他们的身份。用户可以灵活地为与其简档(例如,年龄、性别)和状态(例如,情感)相关的不同说话者属性设置匿名请求。我们不会默认匿名所有个人信息,因为根据匿名音频的应用,某些属性需要以未经修改的形式保留。特别是由于智能设备的用户往往不知道他们的隐私受到保护或侵犯,我们专注于为用户提供尽可能多的可控性和透明度,同时保持可用性和有效性。此外,为了将隐私支持扩展到非英语使用者,我们在建议的系统中包括一个多语言开关,该开关选择与语言相关的组件并向多语言组件通知有关输入语言的信息。在这项建议中,我们将重点放在德国使用的主要语言(官方或外国),但提出一种可扩展到其他语言的方法。
英文摘要
Automatic speech processing enables useful applications, like speech assistants or automatic transcription services. However, a speech signal contains more data than is usually needed for the task, such as paralinguistic information about the speaker. As classification models with objectives like speaker identification or speech emotion recognition become advanced, this information poses serious privacy threats, often without the speaker’s knowledge. Service providers might abuse the data for speaker profiling, or it might fall into the hands of attackers during data transmission or storage. Therefore, anonymizing the audio - to preserve voice privacy - immediately after recording and before further processing has increased its relevance and importance, especially after the introduction of the General Data Protection Regulation (GDPR) of the European Union. The idea is to modify speech recordings such that the link between the original speaker and the audio is destroyed, for instance, by using the voice of a different speaker. In theory, if the anonymization has been successful, no further actions have to be taken to protect other personal attributes in the data, like speaker traits (e.g., gender, ethnic origin) or speaker states (e.g., health state, emotions). In this project, we propose a voice privacy framework that gives users control over which personal information, including but not limited to their identity, they want to protect in their speech before sharing it with an external service. The user can flexibly set anonymization requests for different speaker attributes related to their profile (e.g., age, gender) and state (e.g., emotion). We will not anonymize all personal information by default because, depending on the application of the anonymized audio, some attributes need to be preserved in an unmodified form. Especially since users of smart devices often do not know about the protection or violation of their privacy, we focus on providing as much controllability and transparency to the user as possible while keeping usability and effectiveness. Furthermore, to extend privacy support to non-English speakers, we include a multilingual switch in our proposed system that selects language-dependent components and informs multilingual components about the input language. In this proposal, we will focus on the major languages spoken in Germany (official or foreign) but propose a method that is extendable to other languages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金