AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models

AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
复制标题

DOI:
10.1145/3582700.3582722
复制
发表时间:
2023-03
期刊:
Proceedings of the Augmented Humans International Conference 2023
影响因子:
--
通讯作者:
Kazuki Kawamura;J. Rekimoto
Kazuki Kawamura;J. Rekimoto
中科院分区:
其他
文献类型:
--
作者:
Kazuki Kawamura;J. Rekimoto

文献摘要

相似文献

由于人类收听音频和观看视频的速度比实际观察到的速度更快,因此我们经常以更高的播放速度收听或观看这些内容,以提高内容理解的时间效率。为了进一步利用这种能力,人们开发了根据用户状况和内容类型自动调整播放速度的系统,以帮助更有效地理解时间序列内容。然而,这些系统仍有空间进一步扩展人类的快速聆听能力,通过生成针对更精细的时间单位优化的播放速度的语音并将其提供给人类。在这项研究中,我们确定人类是否可以听到优化后的语音,并提出了一种系统,该系统可以以小至音素的单位自动调整播放速度,同时确保语音清晰度。该系统使用语音识别器分数作为人类听到特定语音单元的程度的代理,并将语音播放速度最大化到人类可以听到的程度。这种方法可用于产生快速但易懂的语音。在评估实验中,我们在盲测中比较了恒定快速播放的语音和该方法生成的灵活加速的语音,并证实该方法产生的语音更容易听。
Since humans can listen to audio and watch videos at faster speeds than actually observed, we often listen to or watch these pieces of content at higher playback speeds to increase the time efficiency of content comprehension. To further utilize this capability, systems that automatically adjust the playback speed according to the user’s condition and the type of content to assist in more efficient comprehension of time-series content have been developed. However, there is still room for these systems to further extend human speed-listening ability by generating speech with playback speed optimized for even finer time units and providing it to humans. In this study, we determine whether humans can hear the optimized speech and propose a system that automatically adjusts playback speed at units as small as phonemes while ensuring speech intelligibility. The system uses the speech recognizer score as a proxy for how well a human can hear a certain unit of speech and maximizes the speech playback speed to the extent that a human can hear. This method can be used to produce fast but intelligible speech. In the evaluation experiment, we compared the speech played back at a constant fast speed and the flexibly speed-up speech generated by the proposed method in a blind test and confirmed that the proposed method produced speech that was easier to listen to.