How to teach machines to read human rights reports and identify judgments at scale

How to teach machines to read human rights reports and identify judgments at scale
复制标题

DOI:
10.1080/14754835.2019.1671174
复制
发表时间:
2020-01
影响因子:
1.9
通讯作者:
Baekkwan Park;Kevin Greene;Michael Colaresi
Baekkwan Park;Kevin Greene;Michael Colaresi
中科院分区:
法学3区
文献类型:
--
作者:
Baekkwan Park;Kevin Greene;Michael Colaresi

文献摘要

相似文献

摘要大赦国际、人权观察和美国国务院等人权监测机构提供的信息日益增多,为以更高的分辨率衡量镇压和人权保护提供了新的机会。然而,到目前为止,大多数试图自动构建文本报告的方法都使用简单、低维的观察,例如忽略语法和词序的单词计数。虽然这些陈述对某些应用很有用,但它们限制了学者和政策制定者可以从人权报告中提取的推论。在这篇文章中,我们提出了一个新的系统,脉冲星,它考虑了句法和词序。这个系统独一无二地允许研究人员从文本中提取判决和被判决的方面/权利。我们说明,这些更详细的信息对于改善对身体完整权利和妇女政治权利的预测是有用的,而且还有助于生成比传统规范更具解释性的机器学习模型。后一项好处有望将人权文本的定性和定量分析连贯地联系起来。
Abstract The accelerating availability of information from human rights monitors such as Amnesty International, Human Rights Watch, and the US State Department has led to new opportunities to measure repression and human rights protections in higher resolution. However, to date, most approaches that attempt to automatically structure textual reports use simple, lower-dimensional observations such as the counts of words that ignore syntax and word order. While these representations are useful for some applications, they limit the inferences scholars and policy-makers can extract from human rights reports. In this article, we present a new system, PULSAR, that takes syntax and word order into account. This system uniquely allows researchers to extract both the judgements and the aspects/rights being judged from texts at scale. We illustrate that this more detailed information is useful both for improving predictions of physical integrity rights and women's political rights, but also for generating machine learning models that are more interpretable than conventional specifications. This latter benefit holds the promise of coherently connecting qualitative and quantitative analyses of human rights texts.