Silent Paralinguistics
Silent Paralinguistics
批准号:
514018165
负责人:
Professor Dr.-Ing. Björn Schuller
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
语言是人类与生俱来的能力,也是使我们成为社会物种的核心部分。将说话者的嘴唇藏在口罩后面会降低听众的表现和信心,同时增加感知上的努力。听力受损和非母语人士面临着更大的挑战。除了这些问题,面具还阻碍了人际交流的副语言学,即说话的方式。对于声学语音,可以使用计算副语言方法自动识别副语言。静音语音接口(SSI)即使在声音信号严重退化或不可用的情况下也能实现语音通信。SSI的目标是为无声的说话者产生语音,否则就会从语音产生过程本身产生的生物信号中使个人静音。这种与语音相关的生物信号包括来自发音器、发音肌肉活动、神经通路和大脑本身的信号。表面肌电信号(EMG)捕捉关节肌肉的活动,已成功地应用于SSIS。使用基于EMG的SSI,无声语音被转换为文本或直接转换为可听语音。尽管取得了重大进展,但副语言学的缺乏仍然是SSI用户的一个主要问题。在该方案中,我们将无声语音接口与计算副语言学相结合,为“无声副语言学”奠定了基础。SP的目标是首先从无声言语产生过程中与语音相关的生物信号中推断说话人的状态和特征,然后将推断出的副语言信息用于更自然的基于SSI的口语会话。我们将研究礼貌和挫折感作为说话人的状态,以及身份和个性作为说话人的特征。作为开发SP方法的基础,我们将记录和标记来自100名参与者的数据,我们将通过包括游戏场景和通过添加令人愤怒的游戏元素来诱导他们的礼貌演讲。基于这些数据,我们将调查从无声产生的语音的EMG信号中能够很好地预测说话人的状态和特征。为此,我们将研究和比较两种方法:直接SP和间接SP,前者直接根据EMG特征预测特征和状态,后者先将EMG转换为声学特征,然后根据声学特征预测特征和状态。此外,我们将优化副语言预测在SSI中的集成,以生成最合适的声学信号。用于多说话人EMG到语音转换的深层生成模型将以特征和状态预测为条件,以便产生的声学信号反映预期的情感意义。建立了一个EMG-SSI原型,最终验证了SP增强的声学语音信号是否在自然度和用户接受度方面改善了口语通信的可用性。
英文摘要
Speech is a natural human ability and a core part of what makes us a social species. Concealing the speaker’s lips behind a face mask lowers listener performance and confidence, while increasing perceptual effort. Hearing impaired and non-native-speakers face even greater challenges. Beyond these issues, masks impede the paralinguistics of interpersonal communication, i.e. the way something was said. For acoustic speech, paralinguistics can be automatically recognized using Computational Paralinguistics methods. Silent Speech Interfaces (SSIs) enable spoken communication even when the acoustic signal is severely degraded or unavailable. SSIs aim to generate speech for silent speakers and otherwise mute individuals from biosignals that result from the speech production process itself. Such speech-related biosignals encompass signals from the articulators, articulatory muscle activity, neural pathways and the brain itself. Surface electromyography (EMG), which captures the activity of the articulatory muscles has been successfully applied to SSIs. With EMG-based SSIs, silently spoken speech is converted into text or directly into audible speech. Despite major advances, the lack of paralinguistics remains a major issue for SSI users. In this proposal, we combine Silent Speech Interfaces with Computational Paralinguistics to lay the foundation for “Silent Paralinguistics (SP)”. SP aims to firstly infer speaker states and traits from speech-related biosignals during silent speech production and secondly to use this inferred paralinguistic information for a more natural SSI-based spoken conversation. We will study politeness and frustration as speaker states, as well as identity and personality as speaker traits. As basis for the development of SP methods, we will record and label data from 100 participants, from whom we will elicit polite speech by including game scenarios and frustration by adding infuriating game elements. Based on these data, we will investigate how well speaker states and traits can be predicted from EMG signals of silently produced speech. To this end, we will study and compare two approaches: direct SP, which predicts traits and states directly from the EMG features, and indirect SP, which first converts EMG to acoustic features and then predicts traits and states from the acoustic features. Furthermore, we will optimize the integration of paralinguistic predictions in SSI to generate the most appropriate acoustic signals. Deep generative models for multi-speaker EMG-to-speech conversion will be conditioned on traits and state predictions, such that the produced acoustic signals reflect the intended affective meaning. An EMG-SSI prototype is established to finally validate whether the SP-enhanced acoustic speech signal improves the usability of spoken communication in terms of naturalness and user acceptance.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Kontextsensitive automatische Erkennung spontaner Sprache mit BLSTM-Netzwerken
-
批准号:193507010
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2011
-
负责人:Professor Dr.-Ing. Björn Schuller
-
依托单位:
Nichtnegative Matrix-Faktorisierung zur störrobusten Merkmalsextraktion in der Sprachverarbeitung
-
批准号:168309859
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:2010
-
负责人:Professor Dr.-Ing. Björn Schuller
-
依托单位:
Agent-based Unsupervised Deep Interactive 0-shot-learning Networks Optimising Machines' Ontological Understanding of Sound (AUDI0NOMOUS)
-
批准号:442218748
-
项目类别:Reinhart Koselleck Projects
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Professor Dr.-Ing. Björn Schuller
-
依托单位: