课题基金 / 基金详情

项目摘要

项目成果

Antonia David Vitela的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):语音通信的基本挑战之一是语音产生/声学的可变性。说话的人在声道的大小和形状、方言和说话习惯上各不相同。这些差异都会影响声学输出。尽管声学信号缺乏不变性,但听众可以正确地感知许多不同说话者的语音。这种使人的感知适应说话者的特定声学结构的能力已经被研究了五十多年。对这种现象的普遍解释是,听者构建了特定于谈话的表征,这些表征可以作为后续语音的参考。具体地,认为听众可以创建声学和音素之间的映射,或者提取每个单独说话者的声道解剖结构和形状。拟议中的研究集中在另一种解释,这需要一个更一般的听觉方法。来自先前研究的数据表明,听众可能正在计算说话者语音的平均频谱表示(长期平均频谱- LTAS),并将其用作所指。该过程/表示不是语音特定的,但仍然可以适应一些说话者特定的可变性。在以前的工作中,我已经开发了一个依赖于LTAS计算的感知适应模型。该项目的目标是通过确定对LTAS有效表示的更准确估计并将其预测与感知学习方法的预测进行比较来进一步开发和测试该模型。为了实现这些目标,必须确定监听器计算LTAS的时间窗口(目标#1)。该项目包括一系列的实验中,前面的上下文被添加到一个目标声音,以确定对目标的分类效果。通过增加上下文的持续时间(并让每个持续时间改变LTAS),我可以确定上下文在多大程度上有效地引发感知效果。这些研究的创新之一是,语音合成使用一个现实的声道模型,允许由现实的发音约束的声学控制。它还允许我创建不同的“说话者”,具有可知的解剖/发音差异。这个模型将与传统的方法进行测试,以预测听众接触新的“方言”或“口音”的影响。元音的产生将被转移,以产生可学习的元音分类的差异,为听众,但这些变化将有独立的影响,对说话人的LTAS。通过这种方式,我将能够测试哪个模型最好地解释了感知数据。这种模型的发展将界定听众适应解剖学差异、口音、方言甚至运动性语言障碍的能力。它还提供了信号中哪些信息对于自适应复杂声音感知是重要的指示,这些信息可能会因助听器和耳蜗植入物中的信号处理而失真。 公共卫生相关性:这项研究将提供深入了解的过程/代表参与的能力,听众,以适应语音的变化所产生的差异,包括解剖,说话风格,口音和运动障碍的谈话者的特点。目前的助听器和人工耳蜗系统可能会破坏其中一些信息,而这些信息对于强大的语音感知至关重要。研究结果可能会影响这些听力设备的未来发展,以及提高可懂度的策略。
英文摘要
DESCRIPTION (provided by applicant): One of the fundamental challenges for communication by speech is the variability in speech production/acoustics. Talkers vary in the size and shape of their vocal tract, in dialect, and in speaking mannerisms. These differences all impact the acoustic output. Despite this lack of invariance in the acoustic signal, listeners can correctly perceive the speech of many different talkers. This ability to adapt one's perception to the particular acoustic structure of a talker has been investigated for over fifty years. The prevailing explanation for this phenomenon is that listeners construct talk-specific representations that can serve as referents for subsequent speech sounds. Specifically, it is thought that listeners may either be creating mappings between acoustics and phonemes or extracting the vocal tract anatomy and shape for each individual talker. The proposed research focuses on an alternative explanation, which takes a more general auditory approach. Data from previous studies has indicated that listeners may be calculating an average spectral representation (long term average spectrum - LTAS) of a talker's speech and using that as a referent. This process/representation is not speech-specific but can still accommodate some of the talker-specific variability. In previous work, I have developed a model of perceptual adaptation that relies on the computation of the LTAS. The goal of this project is to further develop and test this model by determining a more accurate estimate of the effective representation of the LTAS and comparing its predictions to those of perceptual learning approaches. In order to accomplish these goals, the time window over which the LTAS is computed by listeners must be determined (Aim #1). The project includes a series of experiments in which preceding context is added to a target sound to determine the effect on categorization of the target. By increasing the duration of the context (and having each duration change the LTAS), I can determine how much of the context is effective in eliciting a perceptual effect. One of the innovations of these studies is that the speech is synthesized using a realistic vocal tract model, allowing acoustic control constrained by realistic articulations. It also allows me to create different "talkers" with knowable anatomical/articulatory differences. This model will be tested against a traditional approach in predicting the effect on listeners of being exposed to novel "dialects" or "accents". Vowel productions will be shifted to produce learnable differences in vowel categorization for listeners, but these shifts will have independent effects on the LTAS of the talker. In this way, I will be able to test which model best explains the perceptual data. The development of such a model will delimit the ability of listeners to accommodate variations due to anatomical differences, accent, dialect and even motor speech disorders. It also provides an indication of what information is important in the signal for adaptive complex sound perception that may be distorted by signal processing in hearing aids and cochlear implants. PUBLIC HEALTH RELEVANCE: This research will provide insight into the processes/representations involved in the ability of listeners to accommodate the variability in speech arising from differences in talker characteristics including anatomy, speaking style, accent and motor disability. Current hearing aid and cochlear implant systems can disrupt some of this information that could be critical for robust speech perception. The results could impact the future development of these hearing devices, as well as strategies for improving intelligibility.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
General Auditory Model of Adaptive Perception of Speech
  • 批准号:
    8211404
  • 项目类别:
  • 资助金额:
    $4.22万
  • 财政年份:
    2011
  • 负责人:
    Antonia David Vitela
  • 依托单位:
海外基金