Inferring missing metadata from environmental policy texts

Inferring missing metadata from environmental policy texts
复制标题

DOI:
10.18653/v1/w19-2506
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Steven Bethard;Egoitz Laparra;Sophia Wang;Yiyun Zhao;Ragheb Al-Ghezi;Aaron M. Lien;L. López-Hoffman
Steven Bethard;Egoitz Laparra;Sophia Wang;Yiyun Zhao;Ragheb Al-Ghezi;Aaron M. Lien;L. López-Hoffman
中科院分区:
其他
文献类型:
--
作者:
Steven Bethard;Egoitz Laparra;Sophia Wang;Yiyun Zhao;Ragheb Al-Ghezi;Aaron M. Lien;L. López-Hoffman

文献摘要

相似文献

《国家环境政策法》(NEPA)提供了一个关于美国在过去50年中如何制定环境政策的数据库。遗憾的是,没有关于这一信息的中央数据库,而且信息量太大,无法人工评估。我们描述了我们的努力,使系统的研究,美国环境政策的提取和组织元数据的文本NEPA文件。我们的贡献包括收集超过40,000份与NEPA相关的文件,并评估基于规则的基线,这些基线确定了三项重要任务的难度:确定牵头机构,调整文件版本和检测重复使用的文本。
The National Environmental Policy Act (NEPA) provides a trove of data on how environmental policy decisions have been made in the United States over the last 50 years. Unfortunately, there is no central database for this information and it is too voluminous to assess manually. We describe our efforts to enable systematic research over US environmental policy by extracting and organizing metadata from the text of NEPA documents. Our contributions include collecting more than 40,000 NEPA-related documents, and evaluating rule-based baselines that establish the difficulty of three important tasks: identifying lead agencies, aligning document versions, and detecting reused text.