Auditory Scene Analysis Principles Improve Image Reconstruction Abilities of Novice Vision-to-Audio Sensory Substitution Users.

Auditory Scene Analysis Principles Improve Image Reconstruction Abilities of Novice Vision-to-Audio Sensory Substitution Users.
复制标题

DOI:
10.1109/embc46164.2021.9630296
复制
发表时间:
2021-11
期刊:
Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

“vOICe”等感官替代设备(ssd)通过将视觉高度、亮度和横向度分别转换为听觉音高、音量和平移/时间来保存声音中的视觉信息。然而,用户很难识别或跟踪多个同时呈现的音调——这是区分物体形状的上下边缘所必需的技能。我们探索如何通过使用听觉场景分析(ASA)启发的图像超声来解决这些缺陷。在这里,有视力的受试者(N=25)听不同的音乐经验,然后重建,由同时呈现的上下线组成的复杂形状。使用vOICe对复杂的形状进行声学处理,其中上下线仅在音高上变化(即vOICe的“不变”默认设置),或者使用一条线降级以改变其听觉音色或音量。结果发现,随着受试者先前音乐经验的增加,整体表现也会提高。单因素方差分析显示,发声方式和音乐体验对演奏效果均有显著影响,但两者之间无交互作用。与vOICe的“未改变”音高映射相比,当通过音色或音量调制改变低音线时,受试者具有明显更好的图像重建能力。相比之下,改变上行线只能帮助用户识别未改变的下行线。总之,将ASA原理添加到视觉到音频的固态硬盘中,可以提高受试者的图像重建能力,即使这也会减少与任务相关的全部信息。未来的固态硬盘应该寻求利用这些发现来提高新手用户的能力和使用固态硬盘作为视觉康复工具。
Sensory substitution devices (SSDs) such as the ‘vOICe’ preserve visual information in sound by turning visual height, brightness, and laterality into auditory pitch, volume, and panning/time respectively. However, users have difficulty identifying or tracking multiple simultaneously presented tones – a skill necessary to discriminate the upper and lower edges of object shapes. We explore how these deficits can be addressed by using image-sonifications inspired by auditory scene analysis (ASA). Here, sighted subjects (N=25) of varying musical experience listened to, and then reconstructed, complex shapes consisting of simultaneously presented upper and lower lines. Complex shapes were sonified using the vOICe, with either the upper and lower lines varying only in pitch (i.e. the vOICe’s ‘unaltered’ default settings), or with one line degraded to alter its auditory timbre or volume. Results found that overall performance increased with subjects’ years of prior musical experience. ANOVAs revealed that both sonification style and musical experience significantly affected performance, but with no interaction effect between them. Compared to the vOICe’s ‘unaltered’ pitch-height mapping, subjects had significantly better image-reconstruction abilities when the lower line was altered via timbre or volume-modulation. By contrast, altering the upper line only helped users identify the unaltered lower line. In conclusion, adding ASA principles to vision-to-audio SSDs boosts subjects’ image-reconstruction abilities, even if this also reduces total task-relevant information. Future SSDs should seek to exploit these findings to enhance both novice user abilities and the use of SSDs as visual rehabilitation tools.