The effect of speaker sampling in likelihood ratio based forensic voice comparison

The effect of speaker sampling in likelihood ratio based forensic voice comparison
复制标题

基于似然比的法医语音比较中说话人采样的效果

DOI:
--
复制
发表时间:
2019
影响因子:
0.4
通讯作者:
P. Foulkes
P. Foulkes
中科院分区:
人文科学4区
文献类型:
--
作者:
B. Wang;Vincent Hughes;P. Foulkes

文献摘要

被引文献

相似文献

在法医声音比较(FVC)领域,专家们越来越多的压力,以证明他们在个案工作中得出的结论的有效性和可靠性。一个充分的数据驱动的方法,利用说话人的数据库来计算数值似然比(LR)的好处是,它是可能的估计有效性和可靠性的经验。然而,很少有人知道LR输出的稳定性作为一个函数的特定扬声器采样用于在训练,测试和参考数据集。本研究使用两个大的共振峰数据集:粤语句末助词/a/和英式英语填充停顿UM来解决这个问题。实验重复100次,改变1)训练、测试和参考说话者,2)仅训练说话者,3)仅测试说话者,以及4)仅参考说话者。结果表明,在所有三个集合中改变扬声器对粤语和英语变量的系统稳定性的影响最大,与CLLR变化从0.60到0.97为/a/和0.32到1.33为UM。然而,这种可变性主要是由于测试集中的不确定性的影响。仅改变训练扬声器对/a/的系统稳定性的影响最小(Cllr范围:0.76至0.88),而改变参考扬声器对UM的影响最小(Cllr范围:0.40至0.54)。结果表明,在基于LR的FVC中,重要的是评估作为所使用的扬声器的样本(Cllr范围)的函数的系统的稳定性,而不是仅基于每个集合中的扬声器的一个配置报告单个Cllr值。这项研究有助于报告LR计算中的不确定性的一般性辩论。
Within the field of forensic voice comparison (FVC), there is growing pressure for experts to demonstrate the validity and reliability of the conclusions they reach in casework. One benefit of a fully data-driven approach that utilises databases of speakers to compute numerical likelihood ratios (LRs) is that it is possible to estimate validity and reliability empirically. However, little is known about the stability of LR output as a function of the specific speakers sampled for use in the training, test and reference data sets. The present study addresses this issue using two large sets of formant data: Cantonese sentence final particle /a/ and British English filled pauses UM. Experiments were replicated 100 times varying the 1) training, test and reference speakers, 2) training speakers only, 3) test speakers only, and 4) reference speakers only. The results show that varying the speakers in all three sets has the greatest effect on system stability for both the Cantonese and English variables, with the Cllr varying from 0.60 to 0.97 for /a/ and 0.32 to 1.33 for UM. However, this variability is primarily due to the effects of uncertainty in the test set. Varying only the training speakers has the least effect on system stability for /a/ (Cllr range: 0.76 to 0.88), while varying reference speakers has the smallest effect for UM (Cllr range: 0.40 to 0.54). The results indicate that in LR-based FVC it is important to assess the stability of the system as a function of the samples of speakers used (Cllr range) rather than just reporting a single Cllr value based on one configuration of speakers in each set. The study contributes to the general debate on reporting uncertainty in LR computation.