课题基金 / 基金详情

Using BIG data to understand the BIG picture: Overcoming heterogeneity in speech for forensic applications

Using BIG data to understand the BIG picture: Overcoming heterogeneity in speech for forensic applications
使用大数据了解大局:克服取证应用中语音的异质性
批准号:
ES/N003268/1
负责人:
Erica Gold
金额:
$35.77万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2016
资助国家:
英国
项目状态:
已结题
起止时间:
2016 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
法医言语科学(FSS)是语音学的一个应用分支学科,在涉及语音证据的刑事案件中起着关键作用。在金融监督院内部,法医说话人比对(FSC)涉及对犯罪记录(例如威胁电话)和已知嫌疑人样本(例如警察采访)进行比较。法医语音学专家的作用是就两个样本来自同一说话人的可能性向事实审判官(如法官或陪审团)提出建议。做这样的比较有两个重要的因素。首先,专家将对犯罪记录中的语音特征与犯罪嫌疑人样本的相似度进行评估。其次,专家将评估犯罪样本的相同语音特征在多大程度上可以被认为是特定说话者群体的典型特征。说话的群体通常由年龄、性别和地理区域(或口音)来定义。第二个元素对于提供第一个元素的背景至关重要;嫌疑人的语言可能与犯罪记录中的非常相似,但如果他们表现出与说话者群体相同的语言特征,这可能纯粹是巧合。相反,如果观察到罪犯和嫌疑人的言语特征在他们的说话群体中被认为是非典型的,那么这将为他们是同一个人提供强有力的证据。与FSC相关的一个复杂问题是,用于估计特定说话群体的语音特征是典型的还是非典型的数据,通常称为人口数据,几乎没有可用的数据。人口数据通常是通过收集一组录音来获得的,这些录音包含了年龄、性别和地理区域(或口音)相似的同质群体的声音。不幸的是,收集人口数据所涉及的时间和费用意味着法医语音学家在为案件工作获取此类数据时面临巨大挑战。不同说话者群体之间存在的高度差异使这个问题进一步复杂化。FSS领域的方法学研究表明,确定FSC的正确人口对于准确代表证据的强度至关重要。正是由于这些原因,专家们认为该领域面临的最大问题是人口数据的有限可用性。本研究的主要目的是探索一套新颖的建议方法,以寻求补救上述问题。目前缺乏交换数据的平台,这意味着可能已经收集了某一特定发言者群体的人口数据,而需要这种数据的专家却不知道。该项目旨在通过开发一个共享数据的国际平台,并鼓励同行研究人员和专家参与数据共享,从而结束这种情况。此外,该项目将探讨人口数据可推广到何种程度;具体来说,这将需要确定地理(或地区口音)水平,从而可以定义说话群体。例如,专家可能会将一个人口群体定义为具有利兹口音,而实际上,更普遍地定义为西约克郡的人口就足够了。这显然会对收集人口数据的方式产生影响。为了探索定义人口数据的问题,将收集西约克郡(WY) 200名男性发言者的数据库(包括来自四个城市地区的50名发言者:哈德斯菲尔德,利兹,布拉德福德和韦克菲尔德)。该数据库将用于测试FSC案例使用人口数据的不同口音定义进行模拟时证据强度的敏感性。除了在研究方法上发挥作用外,该数据库本身也可作为个案工作和研究的实用资源。
英文摘要
Forensic speech science (FSS) - an applied sub-discipline of phonetics - has come to play a critical role in criminal cases involving voice evidence. Within FSS, Forensic speaker comparison (FSC) involves the comparison of a criminal recording (e.g. a threatening phone call), and a known suspect sample (e.g. a police interview). It is the role of an expert forensic phonetician to advise the trier of fact (e.g. judge or jury) on the likelihood of the two samples coming from the same speaker. There are two important elements involved in making such a comparison. First, the expert will carry out an assessment of the similarity of the speech characteristics in the criminal recording and the suspect sample. Second, the expert will assess the degree to which the same speech features for the criminal sample can be considered to be typical for a given speaker group. The speaker group will typically be defined by age, sex and geographical region (or accent). This second element is critical in providing context for the first; the suspect could have speech very similar to that in the criminal recording but this could be purely coincidental if they exhibit speech characteristics that are common to their speaker group. In contrast, if the criminal and suspect are observed as having speech features considered as being atypical for their speaker group then this would provide strong evidence for it being the same speaker.One complication associated with FSC is that data to estimate whether a speech feature is typical or atypical for the given speaker group, commonly known as population data, are scarcely available. Population data are typically obtained by collecting a set of recordings containing the voices of a homogeneous group of speakers similar in age, sex, and geographical region (or accent). Unfortunately, the time and expense involved in the collection of population data means that forensic phoneticians face a huge challenge in obtaining such data for casework. This problem is further complicated by the high degree of variation that exists in speech across different speaker groups. Methodological research in the field of FSS has demonstrated that identifying the correct population for a FSC is vital in accurately representing the strength of evidence. It is largely for these reasons that experts argue that the biggest problem facing the field is the limited availability of population data.The primary aim of this research is to explore a novel set of proposed methods that seek to remedy the aforementioned problems. The current lack of a platform on which to exchange data means that population data for a specific speaker group might have already been collected, unbeknown to experts in need of such data. This project intends to bring an end to this type of scenario by developing an international platform on which to share data, and also encouraging fellow researchers and experts to participate in data sharing. In addition, the project will explore the extent to which population data are generalizable; specifically, this will entail identifying the geographical (or regional accent) level at which speaker groups can be defined. For example, an expert might define a population group as having a Leeds accent, when in actuality a population defined more generally as West Yorkshire would suffice. This would clearly have implications for the way in which population data would be collected.In order to explore the issue of defining the population data, a West Yorkshire (WY) database of 200 male speakers will be collected (including 50 speakers from each of the four urban areas: Huddersfield, Leeds, Bradford, and Wakefield). The database will be used to test the sensitivity of the strength of evidence when FSC cases are simulated using varying definitions of accent for the population data. In addition to serving methodological purpose, the WY database will also serve as a practical resource for casework and research in its own right.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
Speaker identification using laughter in a close social network
在紧密的社交网络中使用笑声来识别说话者
DOI: 10.1558/ijsll.34552
发表时间: 2017
期刊: International Journal of Speech Language and the Law
影响因子: 0.4
作者: [Land E]
通讯作者: Land E
Variation in the FACE Vowel across West Yorkshire: Implications for Forensic Speaker Comparisons
西约克郡面部元音的变化:对法医说话者比较的影响
DOI: 10.21437/interspeech.2018-1944
发表时间: 2018
期刊:
影响因子: --
作者: [Earnshaw K]
通讯作者: Earnshaw K
"We don't pronounce our t's around here": Realisations of /t/ in West Yorkshire English
“我们在这里不发音我们的 t”:/t/ 在西约克郡英语中的实现
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者: [Earnshaw K]
通讯作者: Earnshaw K
An introduction to the West Yorkshire Regional English Database (WYRED)
西约克郡地区英语数据库(WYRED)简介
DOI: --
发表时间: 2016
期刊: Transactions of the Yorkshire Dialect Society
影响因子: --
作者: [Gold, E.]
通讯作者: Gold, E.
共 6 条
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
    ARF鸟苷酸交换因子BIG1介导ACSL4依赖性铁死亡在非酒精性脂肪性肝炎中的作用及机制研究
    • 批准号:
      --
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2022
    • 负责人:
      游艳
    • 依托单位:
    基于Big Code深度背景增强的Android应用代码反混淆研究
    • 批准号:
      61972290
    • 项目类别:
      面上项目
    • 资助金额:
      60.0万元
    • 批准年份:
      2019
    • 负责人:
      刘进
    • 依托单位:
    BIG1介导STING囊泡转运在抗肺癌免疫反应中的作用及分子机制
    • 批准号:
      81903639
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      21.0万元
    • 批准年份:
      2019
    • 负责人:
      张素林
    • 依托单位: