Toward Fail-Safe Speaker Recognition: Trial-Based Calibration With a Reject Option

Toward Fail-Safe Speaker Recognition: Trial-Based Calibration With a Reject Option
复制标题

DOI:
10.1109/taslp.2018.2875794
复制
发表时间:
2019-01-01
影响因子:
5.4
通讯作者:
Lawson, Aaron
Lawson, Aaron
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ferrer, Luciana;Nandwana, Mahesh Kumar;Lawson, Aaron

文献摘要

被引文献

相似文献

大多数说话人识别系统的输出分数不能直接解释为独立的值。因此,通常对分数执行校准步骤,将其转换为具有明确概率解释的适当似然比。标准校准方法使用一个线性函数来转换系统分数,该线性函数使用选择的数据来训练,以密切匹配评估条件。但是,当评估条件未知时,这种选择是不可行的。在之前的工作中,我们提出了一种称为基于试验的校准(TBC)的校准方法。TBC使用从候选训练集中动态选择的数据为每个测试试验训练一个单独的校准模型,以匹配试验的条件。在这项工作中,我们扩展了TBC方法,提出了:1)一种新的相似性度量,用于选择训练数据,其结果比原始工作中提出的方法显著提高;2)一个新的选项,当没有足够的匹配数据可用于训练校准模型时,系统可以拒绝试验;3)使用正则化来提高每次试验训练的校准模型的稳健性。我们在由多个条件组成的开发集和联邦调查局多条件说话人识别数据集上对所提出的算法进行了测试,结果表明,当有匹配的校准数据可供选择时,所提出的方法在大多数条件下将校准损失降低到接近0的值,并且可以拒绝大多数无法获得相关校准数据的试验。
The output scores of most of the speaker recognition systems are not directly interpretable as stand-alone values. For this reason, a calibration step is usually performed on the scores to convert them into proper likelihood ratios, which have a clear probabilistic interpretation. The standard calibration approach transforms the system scores using a linear function trained using data selected to closely match the evaluation conditions. This selection, though, is not feasible when the evaluation conditions are unknown. In previous work, we proposed a calibration approach for this scenario called trial-based calibration (TBC). TBC trains a separate calibration model for each test trial using data that is dynamically selected from a candidate training set to match the conditions of the trial. In this work, we extend the TBC method, proposing: 1) a new similarity metric for selecting training data that result in significant gains over the one proposed in the original work; 2) a new option that enables the system to reject a trial when n ot enough matched data are available for training the calibration model; and 3) the use of regularization to improve the robustness of the calibration models trained for each trial. We test the proposed algorithms on a development set composed of several conditions and on the Federal Bureau of Investigation multi-condition speaker recognition dataset, and we demonstrate that the proposed approach reduces calibration loss to values close to 0 for most of the conditions when matched calibration data are available for selection, and that it can reject most of the trials for which relevant calibration data are unavailable.