Humans and machines: novel methods for testing speaker recognition performance
Humans and machines: novel methods for testing speaker recognition performance
批准号:
AH/T012978/1
负责人:
Vincent Hughes
金额:
$25.63万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --
中文摘要
作为人类,我们经常使用声音作为识别人的一种手段-例如,当有人打电话给我们,或者从另一个房间向我们喊叫时。虽然人类相对擅长识别熟悉的声音,或者至少是可预测的声音,但识别不熟悉的声音要困难得多。这往往是法医和调查背景下的任务。在这种情况下,对未知罪犯和已知嫌疑人的声音进行比较,最终目的是评估它们属于同一个人的可能性。在世界各地,越来越多的说话人识别机器(即软件)用于这些目的。然而,一个关键的问题仍然没有答案:机器是否能像人类一样识别说话者?这个问题在文献中得到的关注相对较少。研究这个问题的研究都是小规模的,只是使用总体错误率将人类识别的结果与机器识别的结果进行比较。然而,更重要的是了解一种方法可能优于另一种方法的背景,以及结合这些方法是否有任何好处。在解决这些问题时,我们的研究将更好地了解说话人识别机器的工作原理以及如何改进它们。此外,以前的工作忽略了许多可能影响人类识别性能的因素,如认知偏差。在这个项目中,我们评估了人类判断的可变性,作为不同数量的上下文信息的函数,特别是在刑事审判的背景下,可能有其他与案件相关的信息,甚至是法医专家提供的语音证据,这可能会影响说话人识别任务中涉及的决策过程。为了比较和联合收割机的反应,我们将开发一个定制的电脑游戏,它可以模拟人类的判断,这些判断在概念上与机器产生的判断是等同的。在此过程中,我们还将测试在电脑游戏中使用语音作为中心元素的可行性,这是一个相对较少受到关注的电脑游戏开发领域。该项目有一些具体的研究问题:1.人类和机器在说话人识别方面的表现如何,我们能否通过结合这两种方法来提高性能?那么,这些方法在多大程度上能够获取相同的信息呢?2.在什么情况下(使用具有不同地区口音的说话者和具有不同持续时间和录音质量的不同语音样本),人类的表现优于机器?3.不同的听者群体在说话人比较任务中表现如何?熟悉当地口音能提高成绩吗?4.人类判断在多大程度上受到法医案件中可能出现的背景信息的影响,例如(i)这是一个刑事案件的知识,(ii)案件的其他证据,或(iii)法医专家的意见?
英文摘要
As humans, we regularly use the voice as a means of recognising people - for example, when someone calls us on the telephone, or shouts to us from another room. While humans are relatively good at recognising familiar, or at least predictable voices, identifying unfamiliar voices is much more difficult. This is often the task in forensic and investigative contexts. In such cases, a comparison is made of the voices of an unknown criminal and a known suspect, with the ultimate aim of assessing the likelihood that they belong to the same individual. Increasingly, around the world, speaker recognition machines (i.e. pieces of software) are used for these purposes. However, a critical question remains unanswered: do machines recognise speakers in the way that humans do?This question has received relatively little attention in the literature. The studies that have examined this issue are all small scale and simply compare the results of human recognition with those of machine recognition using overall error rates. However, what is much more important is understanding the contexts in which one method might outperform the other, and whether there is any benefit in combining the approaches. In addressing these issues, our research will provide a better understanding of how speaker recognition machines work and how they might be improved. Further, previous work has overlooked the many factors that may affect human recognition performance, such as cognitive bias. In this project, we assess the variability in human judgements as a function of different amounts of contextual information, especially in the context of a criminal trial where there may be other information pertinent to the case or even a forensic expert providing voice evidence which could influence the decision-making process involved in the speaker recognition task.In order to compare and combine human and machine responses, we will develop a bespoke computer game that elicits human judgements that are conceptually equivalent to those produced by the machine. In doing so, we will also test the viability of using the voice as the central element in a computer game; an area of computer game development that has received relatively little attention.The project has a number of specific research questions:1. How do humans and machines perform at speaker recognition relative to each other, and can we improve performance by combining the two approaches? To what extent, therefore, do these methods capture the same information?2. In what contexts (using speakers with different regional accents and diverse speech samples with varying durations and recording quality) do humans outperform machines?3. How do different listener groups perform in speaker comparison tasks? Does familiarity with the regional accent improve performance? 4. To what extent are human judgements affected by contextual information that may occur in a forensic case, such as (i) the knowledge that it is a criminal case, (ii) other evidence from the case, or (iii) a forensic expert's opinion?
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Eliciting and evaluating likelihood ratios for speaker recognition by human listeners under forensically realistic channel-mismatched conditions
在取证现实的通道不匹配条件下,得出并评估人类听众识别说话人的似然比
DOI:
10.21437/interspeech.2022-490
发表时间:
2022
期刊:
影响因子:
--
作者:
[Hughes V]
通讯作者:
Hughes V
Person-specific automatic speaker recognition: understanding the behaviour of individual speakers for applications of ASR
-
批准号:ES/W001241/1
-
项目类别:Research Grant
-
资助金额:$103.22万
-
财政年份:2022
-
负责人:Vincent Hughes
-
依托单位:
国内基金
海外基金
基于Support Vector Machines(SVMs)算法的智能型期权定价模型的研究
-
批准号:70501008
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2005
-
负责人:曹丽娟
-
依托单位: