Question Answering for Privacy Policies: Combining Computational and Legal Perspectives

Question Answering for Privacy Policies: Combining Computational and Legal Perspectives
复制标题

DOI:
10.18653/v1/d19-1500
复制
发表时间:
2019-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Abhilasha Ravichander;A. Black;Shomir Wilson;Thomas B. Norton;N. Sadeh
Abhilasha Ravichander;A. Black;Shomir Wilson;Thomas B. Norton;N. Sadeh
中科院分区:
其他
文献类型:
--
作者:
Abhilasha Ravichander;A. Black;Shomir Wilson;Thomas B. Norton;N. Sadeh

文献摘要

相似文献

隐私政策是冗长而复杂的文档,用户很难阅读和理解。然而,它们对如何收集、管理和使用用户数据具有法律效力。理想情况下,我们希望使用户能够了解对他们来说重要的问题,并使他们能够有选择地探索这些问题。我们提供了一个由1750个关于移动应用程序隐私策略的问题和3500多个相关答案的专家注释组成的语料库PrivocyQA。我们观察到,强大的神经基线在PrivacyQA上的表现比人类表现低近0.3F1,这表明未来系统有相当大的改进空间。此外,我们使用这个数据集来明确地识别问题可应答性的挑战,这对任何问题回答系统都具有领域一般性的影响。PrivacyQA语料库为问题回答提供了一个具有挑战性的语料库,具有真正的现实实用价值。
Privacy policies are long and complex documents that are difficult for users to read and understand. Yet, they have legal effects on how user data can be collected, managed and used. Ideally, we would like to empower users to inform themselves about the issues that matter to them, and enable them to selectively explore these issues. We present PrivacyQA, a corpus consisting of 1750 questions about the privacy policies of mobile applications, and over 3500 expert annotations of relevant answers. We observe that a strong neural baseline underperforms human performance by almost 0.3 F1 on PrivacyQA, suggesting considerable room for improvement for future systems. Further, we use this dataset to categorically identify challenges to question answerability, with domain-general implications for any question answering system. The PrivacyQA corpus offers a challenging corpus for question answering, with genuine real world utility.