PolicyQA: A Reading Comprehension Dataset for Privacy Policies

PolicyQA: A Reading Comprehension Dataset for Privacy Policies
复制标题

PolicyQA:隐私政策的阅读理解数据集

DOI:
10.18653/v1/2020.findings-emnlp.66
复制
发表时间:
2020
期刊:
2023 IEEE International Conference on Big Data (BigData)
影响因子:
--
通讯作者:
Kai
Kai
中科院分区:
--
文献类型:
--
作者:
Wasi Uddin Ahmad;Jianfeng Chi;Yuan Tian;Kai

文献摘要

被引文献

相似文献

隐私政策文件冗长而冗长。问答(QA)系统可以帮助用户找到与他们相关且重要的信息。这一领域的先前研究将QA任务框架为从给定问题的政策文件中检索最相关的文本片段或句子列表。相反,我们认为,为用户提供一个短的文本跨度从政策文件减少了从一个冗长的文本段搜索目标信息的负担。在本文中,我们提出了PolicyQA,一个包含25,017个阅读理解风格示例的数据集,这些示例来自115个网站隐私政策的现有语料库。PolicyQA提供了714个人工注释的问题,用于广泛的隐私实践。我们评估了两个现有的神经QA模型,并进行了严格的分析,以揭示PolicyQA提供的优势和挑战。
Privacy policy documents are long and verbose. A question answering (QA) system can assist users in finding the information that is relevant and important to them. Prior studies in this domain frame the QA task as retrieving the most relevant text segment or a list of sentences from the policy document given a question. On the contrary, we argue that providing users with a short text span from policy documents reduces the burden of searching the target information from a lengthy text segment. In this paper, we present PolicyQA, a dataset that contains 25,017 reading comprehension style examples curated from an existing corpus of 115 website privacy policies. PolicyQA provides 714 human-annotated questions written for a wide range of privacy practices. We evaluate two existing neural QA models and perform rigorous analysis to reveal the advantages and challenges offered by PolicyQA.