Human Rights Texts: Converting Human Rights Primary Source Documents into Data.

Human Rights Texts: Converting Human Rights Primary Source Documents into Data.
复制标题

DOI:
10.1371/journal.pone.0138935
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Tsai M
Tsai M
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Fariss CJ;Linder FJ;Jones ZM;Crabtree CD;Biek MA;Ross AS;Kaur T;Tsai M

文献摘要

被引文献

相似文献

我们推出并公开提供大量数字化的主要来源人权文件,这些文件每年由包括大赦国际、人权观察、人权律师委员会和美国国务院在内的监督机构出版。除了数字化文本之外,我们还提供并描述了文件术语矩阵,这是一种数据集,它系统地组织了人权文件语料库中每个独特文件的每个独特术语的字数。为了说明这个语料库的重要性,我们描述了在人权界的编码程序的发展和几个现有的分类指标,这些指标是通过对语料库中的人权文件进行人类编码而创建的。然后,我们讨论了新的人权语料库和现有的人权数据集如何与各种统计分析和机器学习算法一起使用,以帮助学者了解人权实践和报告如何随着时间的推移而演变。最后,我们讨论了数据集维护、更新和可用性的计划。
We introduce and make publicly available a large corpus of digitized primary source human rights documents which are published annually by monitoring agencies that include Amnesty International, Human Rights Watch, the Lawyers Committee for Human Rights, and the United States Department of State. In addition to the digitized text, we also make available and describe document-term matrices, which are datasets that systematically organize the word counts from each unique document by each unique term within the corpus of human rights documents. To contextualize the importance of this corpus, we describe the development of coding procedures in the human rights community and several existing categorical indicators that have been created by human coding of the human rights documents contained in the corpus. We then discuss how the new human rights corpus and the existing human rights datasets can be used with a variety of statistical analyses and machine learning algorithms to help scholars understand how human rights practices and reporting have evolved over time. We close with a discussion of our plans for dataset maintenance, updating, and availability.