VOCALISE : A forensic automatic speaker recognition system supporting spectral , phonetic , and user-provided features
VOCALISE : A forensic automatic speaker recognition system supporting spectral , phonetic , and user-provided features
复制标题
VOCALISE:法庭自动说话人识别系统,支持频谱、语音和用户提供的功能
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Finnian Kelly
中科院分区:
文献类型:
--
作者:
A. Alexander;Oscar Forth;Alankar Atreya;Finnian Kelly
In this article we present the latest version of VOCALISE (Voice Comparison and Analysis of the Likelihood of Speech Evidence), a forensic automatic system for speaker recognition. VOCALISE, with selectable stateof-the-art and legacy speaker modelling algorithms allows the forensic practitioner to work with spectral features (such as Mel Frequency Cepstral Coefficients (MFCCs)), phonetic features (such as formants), or features of their own choice (such as voice quality metrics, articulation rate, etc.). It is capable of comparing features from a test audio file of a target speaker against features from an audio file of a suspected speaker, or an entire list of suspected speakers, and produces a likelihood score or likelihood ratio for each comparison. It is built with an ‘open-box’ architecture that transparently allows the user to provide their own data to train the system’s algorithms. These algorithms include Gaussian Mixture Modelling (GMM) with (and without) MAP (maximum a posteriori) adaptation, i-vector extraction with PLDA (Probabilistic Linear Discriminant Analysis) and cosine distance comparison. VOCALISE seeks to form a bridge between traditional forensic phonetics-based speaker recognition and forensic automatic speaker recognition.