Vocal Imitations of Non-Vocal Sounds

Vocal Imitations of Non-Vocal Sounds
复制标题

非声音的声音模仿

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
P. Susini
P. Susini
中科院分区:
综合性期刊3区
文献类型:
--
作者:
G. Lemaitre;Olivier Houix;Frédéric Voisin;N. Misdariis;P. Susini

文献摘要

被引文献

相似文献

模仿行为在人类中很普遍,特别是当两个人交流和互动时。口语中的几个表征(拟声词、意音词和音素)也在单词的声音和它所指的东西之间显示出不同程度的象似性。因此,人类说话者在交流声音时使用大量模仿性的发声和手势可能并不奇怪,因为声音特别难以描述。更令人惊讶的是,对日常生活中的非声音(例如汽车驶过的声音)的声音模仿在实践中非常有效:听众用声音模仿比用语言描述更好地识别声音,尽管事实上声音模仿是由特定机械系统(例如汽车行驶)通过不同系统(语音装置)产生的声音的不准确再现。本研究通过实验量化了听者将声音与类别标签匹配的能力,探讨了声音模仿引起的语义表征。该实验使用了三种不同类型的声音:易于识别的声音(人类行为和制造产品的声音),人类声音模仿和计算“听觉草图”(由算法计算创建)。结果表明,最好的声音模仿的表现是相似的最好的听觉草图的声音类别,甚至在某些情况下,所指的声音本身。更详细的分析表明,声音模仿和所指声音之间的声学距离不足以解释这种性能。分析表明,声乐模仿不是试图尽可能准确地再现所指的声音,而是专注于一些重要的特征,这些特征取决于每个特定的声音类别。这些结果提供了理解人类听众如何存储和访问长期的声音表示的观点,并设置基于发声的人机界面的发展阶段。
Imitative behaviors are widespread in humans, in particular whenever two persons communicate and interact. Several tokens of spoken languages (onomatopoeias, ideophones, and phonesthemes) also display different degrees of iconicity between the sound of a word and what it refers to. Thus, it probably comes at no surprise that human speakers use a lot of imitative vocalizations and gestures when they communicate about sounds, as sounds are notably difficult to describe. What is more surprising is that vocal imitations of non-vocal everyday sounds (e.g. the sound of a car passing by) are in practice very effective: listeners identify sounds better with vocal imitations than with verbal descriptions, despite the fact that vocal imitations are inaccurate reproductions of a sound created by a particular mechanical system (e.g. a car driving by) through a different system (the voice apparatus). The present study investigated the semantic representations evoked by vocal imitations of sounds by experimentally quantifying how well listeners could match sounds to category labels. The experiment used three different types of sounds: recordings of easily identifiable sounds (sounds of human actions and manufactured products), human vocal imitations, and computational “auditory sketches” (created by algorithmic computations). The results show that performance with the best vocal imitations was similar to the best auditory sketches for most categories of sounds, and even to the referent sounds themselves in some cases. More detailed analyses showed that the acoustic distance between a vocal imitation and a referent sound is not sufficient to account for such performance. Analyses suggested that instead of trying to reproduce the referent sound as accurately as vocally possible, vocal imitations focus on a few important features, which depend on each particular sound category. These results offer perspectives for understanding how human listeners store and access long-term sound representations, and sets the stage for the development of human-computer interfaces based on vocalizations.