VOCALISE : A forensic automatic speaker recognition system supporting spectral , phonetic , and user-provided features

VOCALISE : A forensic automatic speaker recognition system supporting spectral , phonetic , and user-provided features
复制标题

VOCALISE:法庭自动说话人识别系统,支持频谱、语音和用户提供的功能

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Finnian Kelly
Finnian Kelly
中科院分区:
--
文献类型:
--
作者:
A. Alexander;Oscar Forth;Alankar Atreya;Finnian Kelly

文献摘要

被引文献

相似文献

在这篇文章中,我们提出了最新版本的VOCALISE(语音比较和分析的语音证据的可能性),一个法医自动系统的说话人识别。VOCALISE具有可选择的最先进的和传统的说话人建模算法,允许法医从业者使用频谱特征(例如梅尔倒谱系数(MFCC))、语音特征(例如共振峰)或他们自己选择的特征(例如语音质量指标、清晰度等)。它能够将来自目标说话者的测试音频文件的特征与来自可疑说话者的音频文件的特征或可疑说话者的整个列表进行比较,并且为每个比较产生似然分数或似然比。它采用“开放式”架构,透明地允许用户提供自己的数据来训练系统的算法。这些算法包括高斯混合模型(GMM)与(或不与)MAP(最大后验)适应,i-向量提取PLDA(概率线性判别分析)和余弦距离比较。VOCALISE试图在传统的基于语音学的说话人识别和法医自动说话人识别之间建立一座桥梁。
In this article we present the latest version of VOCALISE (Voice Comparison and Analysis of the Likelihood of Speech Evidence), a forensic automatic system for speaker recognition. VOCALISE, with selectable stateof-the-art and legacy speaker modelling algorithms allows the forensic practitioner to work with spectral features (such as Mel Frequency Cepstral Coefficients (MFCCs)), phonetic features (such as formants), or features of their own choice (such as voice quality metrics, articulation rate, etc.). It is capable of comparing features from a test audio file of a target speaker against features from an audio file of a suspected speaker, or an entire list of suspected speakers, and produces a likelihood score or likelihood ratio for each comparison. It is built with an ‘open-box’ architecture that transparently allows the user to provide their own data to train the system’s algorithms. These algorithms include Gaussian Mixture Modelling (GMM) with (and without) MAP (maximum a posteriori) adaptation, i-vector extraction with PLDA (Probabilistic Linear Discriminant Analysis) and cosine distance comparison. VOCALISE seeks to form a bridge between traditional forensic phonetics-based speaker recognition and forensic automatic speaker recognition.