课题基金 / 基金详情

Exploring the Potential of Natural Language Processing Techniques in Criminal Justice Agencies

Exploring the Potential of Natural Language Processing Techniques in Criminal Justice Agencies
探索自然语言处理技术在刑事司法机构中的潜力
批准号:
2443762
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
全标题:探索自然语言处理技术在刑事司法机构中的潜力:对假释委员会释放决定中的种族差异的调查拉米审查(2017)强调了在获取统计数据以探索英格兰和威尔士刑事司法(CJ)系统中的种族差异方面的重大困难,并敦促CJ机构解决这一问题。在某些情况下,解决方案只是发布现有数据集的匿名版本。在许多其他情况下,所需数据并不存在。采用新的数据收集程序往往是不切实际的,因为这需要过去十年预算耗尽的机构进行额外投资。为应对这些限制,这一跨学科项目将探索另一种具有成本效益的方法,通过将自然语言处理技术应用于储存在现有司法行政记录中的大量自由文本数据,生成能够突出司法系统潜在差异的新数据。所有终审法院机构都产生大量记录,记录所处理案件的特点和通过的决定。使用内容分析将相关文本信息编码到统计数据集中以手动处理这类记录的传统由来已久,从而能够进行定量分析(例如,Hood,1992;Myers和Talarico,1987)。这些技术的关键问题在于其可伸缩性。手工处理记录的成本/时间与要处理的案件数量成正比,这使得样本要么太少,要么太贵(Pina-Sánchez等人,在出版社)。随着数据科学和人工智能领域的进步,我们提出使用文本挖掘技术来自动进行这样的编码过程。最近,Pina-Sánchez等人。(2019)展示了这种技术在处理判刑记录方面的潜力。然而,这一概念证明仍然受到有关文件有效性和提出的算法的复杂性的重要限制。在之前工作的基础上,并与假释委员会--一个关键的国家CJ机构--合作,该项目将推动这一关键研究领域的方法学前沿。假释委员会进行风险评估,以决定是否可以将囚犯安全释放到社区。在最近的年度调查中,假释委员会(2018a)处理了16,436个“纸质听证会”,通常会记录简短的结构化文本摘要(大约两页长)。这些“听证摘要”记录了案件的主要特征,以及囚犯的人口统计因素,包括他们的种族。项目的目标是:1)开发能够处理“听证摘要”的文本挖掘算法;2)评估这些算法产生的数据的可靠性;3)分析它们产生的数据,以探索假释委员会决策中潜在的种族差异。总之,拟议的跨学科项目及其利用的战略伙伴关系将为博士生提供一个独特的机会,来探索如何在刑事司法结果的背景下将数据科学方法应用于公共利益。通过传播产生的统计数据和方法,该项目还将使法院法官研究的创新思路成为可能,并希望促进法院法官分析员在一系列政府机构中采用新的尖端方法(例如,处理判刑前报告、判决书抄本和几乎任何其他法院法官记录)。这样做,该项目具有非常现实的潜力,可以提供直接的现实潜力,直接回应David Lammy关于数据集的请求,从而揭示CJ系统中任何潜在的差异。
英文摘要
Full title: Exploring the Potential of Natural Language Processing Techniques in Criminal Justice Agencies: An Investigation of Racial Disparities in Release Decisions from the Parole BoardThe Lammy Review (2017) highlights significant difficulties in accessing statistics with which to explore ethnic disparities in the Criminal Justice (CJ) system in England and Wales, and urges CJ agencies to redress this issue. In some instances, the solution simply involves publishing anonymised versions of existing datasets. In many other cases, the required data do not exist. Adopting new data collection processes is often unrealistic since it requires additional investments from agencies that have seen their budgets depleted over the last decade. In response to these constraints, this interdisciplinary project will explore an alternative cost-effective route to generate new data capable of highlighting potential disparities in the CJ system - by applying natural language processing techniques to the large volumes of free-text data stored in existing administrative CJ records. All CJ agencies generate large numbers of records documenting the characteristics of cases processed and decisions adopted. There is a long tradition of manually processing such records using content analysis to code relevant text information into statistical datasets, thus enabling quantitative analyses (e.g. Hood, 1992; Myers and Talarico, 1987). The key problem with these techniques lies in their scalability. The cost/time of processing records manually is directly proportional to the number of cases to be processed, which renders samples either too small or too expensive (Pina-Sánchez et al., In Press). Following advances in the field of Data Science and Artificial Intelligence, we propose the use of text-mining techniques to undertake such coding process automatically. Recently, Pina-Sánchez et al. (2019) demonstrated the potential of such techniques for the processing of sentence records. Yet this proof of concept is still affected by important limitations regarding document validity and the sophistication of algorithms presented. Building on this previous work, and partnering with The Parole Board - a key national CJ agency - the project will push the methodological frontier in this crucial area of research. The Parole Board carries out risk assessments to decide whether prisoners can be safely released into the community. In their latest annual exercise The Parole Board (2018a) processed 16,436 'paper-hearings' for which short structured textual summaries (roughly two-pages long) were routinely recorded. These 'hearing summaries' capture the main characteristics of the case, together with demographic factors of the prisoner including their ethnicity. Analysing a significant sample of these 'hearing summaries' the project objectives are: i) to develop text-mining algorithms capable of processing 'hearing summaries'; ii) to assess the reliability of the data these algorithms generate; and iii) to analyse the data they produce to explore potential racial disparities in Parole Board decisions.In conclusion, the proposed interdisciplinary project and the strategic partnership it leverages will provide the PhD student with a unique opportunity to explore how data science methods can be applied for public good in the context of criminal justice outcomes. By disseminating the statistical data and methods generated, the project will also enable innovative new lines of CJ research, and, it is hoped, facilitate the adoption of new cutting edge methods by CJ analysts across a range of government agencies (e.g. processing of pre-sentence reports, sentence transcripts, and virtually any other CJ records). Doing so the project has a very realistic potential of providing a direct realistic potential of providing a direct response to David Lammy's request for datasets shedding new light on any potential disparities in the CJ system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Transient Receptor Potential 通道 A1在膀胱过度活动症发病机制中的作用
  • 批准号:
    30801141
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    28.0万元
  • 批准年份:
    2008
  • 负责人:
    都书琪
  • 依托单位: