Inferring missing metadata from environmental policy texts
Inferring missing metadata from environmental policy texts
复制标题
DOI:
10.18653/v1/w19-2506
复制
发表时间:
2019-06
期刊:
影响因子:
--
通讯作者:
Steven Bethard;Egoitz Laparra;Sophia Wang;Yiyun Zhao;Ragheb Al-Ghezi;Aaron M. Lien;L. López-Hoffman
中科院分区:
文献类型:
--
作者:
Steven Bethard;Egoitz Laparra;Sophia Wang;Yiyun Zhao;Ragheb Al-Ghezi;Aaron M. Lien;L. López-Hoffman
The National Environmental Policy Act (NEPA) provides a trove of data on how environmental policy decisions have been made in the United States over the last 50 years. Unfortunately, there is no central database for this information and it is too voluminous to assess manually. We describe our efforts to enable systematic research over US environmental policy by extracting and organizing metadata from the text of NEPA documents. Our contributions include collecting more than 40,000 NEPA-related documents, and evaluating rule-based baselines that establish the difficulty of three important tasks: identifying lead agencies, aligning document versions, and detecting reused text.