Mapping Across Feature Spaces in Forensic Voice Comparison: The Contribution of Auditory-Based Voice Quality to (Semi-)Automatic System Testing

Mapping Across Feature Spaces in Forensic Voice Comparison: The Contribution of Auditory-Based Voice Quality to (Semi-)Automatic System Testing
复制标题

取证语音比较中的跨特征空间映射:基于听觉的语音质量对(半)自动系统测试的贡献

DOI:
10.21437/interspeech.2017-1508
复制
发表时间:
2017
期刊:
Biochimica et biophysica acta
影响因子:
--
通讯作者:
Eugenia San Segundo
Eugenia San Segundo
中科院分区:
--
文献类型:
--
作者:
Vincent Hughes;Philip Harrison;P. Foulkes;Peter French;C. Kavanagh;Eugenia San Segundo

文献摘要

被引文献

相似文献

在法医语音比对中,人们越来越关注自动方法和语音方法的结合,以提高法庭语音证据的有效性和可靠性。据此,我们对语音信号的长期测量进行了比较,以评估它们捕获特定说话人的互补信息的程度。使用 MFCC 和(线性和梅尔加权)长期共振峰分布 (LTFD) 进行基于似然比的测试。与基线 MFCC 系统相比,融合自动和半自动系统的性能改进有限,这表明这些措施捕获的特定说话者信息基本相同。性能最佳系统的输出用于评估系统测试中基于听觉的上喉(滤波器)和喉(源)语音质量分析的贡献。结果表明,(半)自动系统中存在问题的说话人在某种程度上可以通过其喉上语音质量特征来预测,最不独特的说话人会产生最弱的证据和最多的错误分类。然而,通过听觉分析仍然可以轻松区分错误分类的对。因此,喉语音质量可能有助于解决(半)自动系统的问题对,从而有可能提高其整体性能。
In forensic voice comparison, there is increasing focus on the integration of automatic and phonetic methods to improve the validity and reliability of voice evidence to the courts. In line with this, we present a comparison of long-term measures of the speech signal to assess the extent to which they capture complementary speaker-specific information. Likelihood ratio-based testing was conducted using MFCCs and (linear and Mel-weighted) long-term formant distributions (LTFDs). Fusing automatic and semi-automatic systems yielded limited improvement in performance over the baseline MFCC system, indicating that these measures capture essentially the same speaker-specific information. The output from the best performing system was used to evaluate the contribution of auditory-based analysis of supralaryngeal (filter) and laryngeal (source) voice quality in system testing. Results suggest that the problematic speakers for the (semi-)automatic system are, to some extent, predictable from their supralaryngeal voice quality profiles, with the least distinctive speakers producing the weakest evidence and most misclassifications. However, the misclassified pairs were still easily differentiated via auditory analysis. Laryngeal voice quality may thus be useful in resolving problematic pairs for (semi-)automatic systems, potentially improving their overall performance.