Speech segmentation and speaker diarisation for transcription and translation

Speech segmentation and speaker diarisation for transcription and translation
复制标题

用于转录和翻译的语音分段和说话人分类

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
M. Sinclair
M. Sinclair
中科院分区:
--
文献类型:
--
作者:
M. Sinclair

文献摘要

被引文献

相似文献

本文概述了与语音分割相关的工作-将音频记录分割成语音和非语音区域,以及扬声器日记-进一步将这些区域分割成与同质扬声器有关的区域。不仅要知道说了什么,还要知道是谁说的,什么时候说的,这有很多有用的应用。除了为语音提供更丰富的转录水平外,我们还将展示这些知识如何提高自动语音识别(ASR)系统的性能,以及如何使下游自然语言处理(NLP)任务(如机器翻译和标点符号恢复)受益。虽然分割和日记化可能看起来是相对简单的任务来描述,但在实践中,我们发现它们非常具有挑战性,并且通常是不明确的问题。因此,我们首先提供了一个形式化的每一个问题的语音声学空间和时间内的细分。在这里,我们看到,当我们想要将这个域划分为我们的目标说话者类别时,任务可能变得非常困难,同时避免驻留在同一空间中的其他类别,例如音素。我们提出了一个理论框架,描述和讨论的任务,以及介绍现有的最先进的方法和研究。目前的Speaker Diarization系统对超参数非常敏感,并且缺乏跨数据集的鲁棒性。因此,我们提出了一种方法,它使用了一系列的甲骨文实验来暴露当前系统的局限性,这些局限性可以归因于系统组件。我们还演示了日志化错误率(DER),在文献中占主导地位的错误度量,是不是一个全面的或可靠的指标的整体性能或错误传播到后续的下游任务。这些结果为我们后续的研究提供了信息。我们发现,作为扬声器日记的先驱,语音分割的任务是系统链中至关重要的第一步。目前的方法通常不考虑口语话语的固有结构。因此,我们探索了一种新的方法,利用话语持续时间之前,以更好地模拟语音的段分布。我们展示了这种方法不仅提高了分割,而且提高了随后的语音识别,机器翻译和扬声器diarization系统的性能。典型的ASR翻译不包括标点符号,用这些信息丰富翻译的任务被称为“标点符号翻译”。其好处不仅是提高了可读性,而且还与NLP系统更好地兼容,这些系统期望类似于传统机器翻译的单位。我们表明
This dissertation outlines work related to Speech Segmentation – segmenting an audio recording into regions of speech and non-speech, and Speaker Diarization – further segmenting those regions into those pertaining to homogeneous speakers. Knowing not only what was said but also who said it and when, has many useful applications. As well as providing a richer level of transcription for speech, we will show how such knowledge can improve Automatic Speech Recognition (ASR) system performance and can also benefit downstream Natural Language Processing (NLP) tasks such as machine translation and punctuation restoration. While segmentation and diarization may appear to be relatively simple tasks to describe, in practise we find that they are very challenging and are, in general, illdefined problems. Therefore, we first provide a formalisation of each of the problems as the sub-division of speech within acoustic space and time. Here, we see that the task can become very difficult when we want to partition this domain into our target classes of speakers, whilst avoiding other classes that reside in the same space, such as phonemes. We present a theoretical framework for describing and discussing the tasks as well as introducing existing state-of-the-art methods and research. Current Speaker Diarization systems are notoriously sensitive to hyper-parameters and lack robustness across datasets. Therefore, we present a method which uses a series of oracle experiments to expose the limitations of current systems and to which system components these limitations can be attributed. We also demonstrate how Diarization Error Rate (DER), the dominant error metric in the literature, is not a comprehensive or reliable indicator of overall performance or of error propagation to subsequent downstream tasks. These results inform our subsequent research. We find that, as a precursor to Speaker Diarization, the task of Speech Segmentation is a crucial first step in the system chain. Current methods typically do not account for the inherent structure of spoken discourse. As such, we explored a novel method which exploits an utterance-duration prior in order to better model the segment distribution of speech. We show how this method improves not only segmentation, but also the performance of subsequent speech recognition, machine translation and speaker diarization systems. Typical ASR transcriptions do not include punctuation and the task of enriching transcriptions with this information is known as ‘punctuation restoration’. The benefit is not only improved readability but also better compatibility with NLP systems that expect sentence-like units such as in conventional machine translation. We show