A review of speaker diarization: Recent advances with deep learning

A review of speaker diarization: Recent advances with deep learning
复制标题

DOI:
10.1016/j.csl.2021.101317
复制
发表时间:
2021-11-20
影响因子:
4.3
通讯作者:
Narayanan, Shrikanth
Narayanan, Shrikanth
中科院分区:
计算机科学3区
文献类型:
--
作者:
Park, Tae Jin;Kanda, Naoyuki;Narayanan, Shrikanth

文献摘要

被引文献

相似文献

说话人日记是一个用与说话人身份相对应的类来标记音频或视频记录的任务,或者简而言之,一个识别“谁在什么时候说话”的任务。在早期,说话人日记算法被开发用于多说话人音频记录的语音识别,以实现说话人自适应处理。随着时间的推移,这些算法作为一个独立的应用程序也获得了自己的价值,为音频检索等下游任务提供特定于说话者的元信息。最近,随着深度学习技术的出现,它推动了语音应用领域研究和实践的革命性变化,扬声器日志化取得了快速进步。在本文中,我们不仅回顾了说话人日记化技术的历史发展,而且还在最近的神经说话人日记化方法的进展。此外,我们还讨论了说话人日志化系统如何与语音识别应用集成,以及最近的深度学习浪潮如何引导这两个组件的联合建模,使其相互补充。通过考虑这些令人兴奋的技术趋势,我们相信本文是对社区的宝贵贡献,通过巩固神经方法的最新发展,从而促进更有效的扬声器日记的进一步进展,提供了一个调查工作。
Speaker diarization is a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify "who spoke when". In the early years, speaker diarization algorithms were developed for speech recognition on multispeaker audio recordings to enable speaker adaptive processing. These algorithms also gained their own value as a standalone application over time to provide speaker-specific metainformation for downstream tasks such as audio retrieval. More recently, with the emergence of deep learning technology, which has driven revolutionary changes in research and practices across speech application domains, rapid advancements have been made for speaker diarization. In this paper, we review not only the historical development of speaker diarization technology but also the recent advancements in neural speaker diarization approaches. Furthermore, we discuss how speaker diarization systems have been integrated with speech recognition applications and how the recent surge of deep learning is leading the way of jointly modeling these two components to be complementary to each other. By considering such exciting technical trends, we believe that this paper is a valuable contribution to the community to provide a survey work by consolidating the recent developments with neural methods and thus facilitating further progress toward a more efficient speaker diarization.