课题基金 / 基金详情

Social Perceptions of Synthetic Speakers

Social Perceptions of Synthetic Speakers
合成扬声器的社会认知
批准号:
423651352
负责人:
Professor Dr.-Ing. Sebastian Möller
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2021-12-31

项目摘要

项目成果

Professor Dr.-Ing. Sebastian Möller的其他基金

相似基金

相关文献

中文摘要
翻译
语音信号自动地在听者中诱导关于说话者的社会认知。通过声学分析和信号处理,已经积累了大量关于社会知觉的相关声学相关性的知识,例如频谱和韵律参数,以及自然语音的感知维度。然而,尽管现代语音合成范例的出现提供了非常高的质量,但仍然不能理解来自自然语音的结果是否也适用于合成语音。因此,主要的研究问题是:合成语音的哪些声学特征会影响对社会说话人特征的主观感知?为了回答这个问题,本项目研究了文本到语音(TTS)合成器在两个潜在应用领域中对两种基本社会属性-能力和仁慈的社会感知:来自医疗保健和客户服务主题的刺激。将结果与早期项目中从自然语音获得的结果进行比较。能力和仁爱是否也是基本的社会属性,或者其他方面是否更相关,这一点得到了检验。对于语音信号,识别了声学参数及其系统学上的相似性和差异性。在方法论层面上,使用最先进的TTS系统创建话语,并在信号层面上系统地修改话语,以便产生用于与人类听众进行经验测试的刺激。在所需的听力和评分测试中采用了众包技术。最终目标是研究如何将声学特征和模式直接结合到现代TTS方法(隐马尔可夫模型、深度神经网络)中,而不是后处理信号处理。这就引出了第二个研究问题:“合成过程的哪些变化会导致说话者的积极感知?”为了达到这一目的,我们采用了现有的说话人转换方法。除了从这项研究中获得的基本知识之外,研究结果将与TTS系统开发人员相关,以便有效地提高特定服务领域的语音质量。
英文摘要
Speech signals automatically induce social perceptions in listeners regarding the speakers. With acoustic analysis and signal manipulation, a great body of knowledge has been accumulated regarding relevant acoustic correlates of social perceptions, such as spectral and prosodic parameters, as well as perceptual dimensions for natural speech. However, despite the advent of modern speech synthesis paradigms providing very high quality, it is yet to be understood, if results from natural speech also hold for synthesized speech. Hence, the major research question is: “Which acoustic features of synthesized speech affect subjective perceptions of social speaker characteristics?”In order to answer this question, this project studies social perception of the two basic social attributions, competence and benevolence, for text-to-speech (TTS) synthesizers in two potential application domains: Stimuli from the topics of healthcare and of customer service. Results are compared to those obtained from natural speech in earlier projects. It is tested whether competence and benevolence also emerge as basic social attributions, or if other dimensions are more relevant. Regarding the speech signal, similarities and differences in acoustic parameters and their systematics are identified. A mid-term result is an acoustic prediction model of the identified social dimensions for synthesized speech.On a methodological level, utterances are created with state-of-the-art TTS systems and systematically modified on the signal level, in order to produce stimuli for empirical testing with human listeners. Crowd-sourcing techniques are applied for the required listening and rating tests. The final goal is to examine, how acoustic features and patterns can be directly incorporated in modern TTS methodologies (Hidden-Markov-Models, Deep Neural Networks) instead of post-processing signal manipulation. This leads to the secondary research question: “Which alterations of the synthesis procedure lead to positive perceptions of speakers?” For this aim, current approaches from speaker conversion are applied.Apart from the fundamental knowledge gained from this research, results will be relevant for TTS system developers, in order to efficiently improve voices for particular service domains.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Quantification of perceived location privacy, and its relationship to privacy behaviour
  • 批准号:
    409241470
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2019
  • 负责人:
    Professor Dr.-Ing. Sebastian Möller
  • 依托单位:
Simulation of Conversation Behavior in Case of Impaired Telephone Transmission
  • 批准号:
    320253669
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2016
  • 负责人:
    Professor Dr.-Ing. Sebastian Möller
  • 依托单位:
Quality Attributes and Overall Quality of Transmitted Speech
Subjective measurement and instrumental estimation of mobile online gaming quality based on perceptual dimensions
  • 批准号:
    279244726
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2015
  • 负责人:
    Professor Dr.-Ing. Sebastian Möller
  • 依托单位:
海外基金