Towards Automatic Classification of Privacy Policy Text

Towards Automatic Classification of Privacy Policy Text
复制标题

迈向隐私政策文本的自动分类

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
N. Sadeh
N. Sadeh
中科院分区:
--
文献类型:
--
作者:
Frederick Liu;Shomir Wilson;Peter Story;Sebastian Zimmeck;N. Sadeh

文献摘要

被引文献

相似文献

隐私政策向互联网用户告知网站、移动的应用程序以及其他产品和服务的隐私惯例。然而,用户很少阅读它们,并且很难理解它们的内容。此外,提供这些政策的实体有时没有动机使它们变得兼容。最近,隐私政策的注释语料库已被引入研究界。它们为机器学习和自然语言处理技术的发展打开了大门,以自动化这些文档的注释。反过来,这些注释可以被传递到接口(例如,Web浏览器插件),帮助用户快速识别和理解相关隐私声明。我们提出了在提取隐私政策段落(本文中称为段)和与专家识别的政策内容类别相关的单个句子方面的进展,使用监督学习的方法。特别是,我们发现相关片段和句子可以分别以0.78和0.66的平均微F1分数进行艾德,比以前的工作有所改进。我们将讨论如何使用本文介绍的技术自动注释约7,000个隐私策略的文本。我们的讨论强调了与我们的分类方法相关的机会和限制。
Privacy policies notify Internet users about the privacy practices of websites, mobile apps, and other products and services. However, users rarely read them and struggle to understand their contents. Also, the entities that provide these policies are sometimes unmotivated to make them com-prehensible. Recently, annotated corpora of privacy policies have been introduced to the research community. They open the door to the development of machine learning and natural language processing techniques to automate the annotation of these documents. In turn, these annotations can be passed on to interfaces (e.g., web browser plugins) that help users quickly identify and understand relevant privacy statements. We present advances in extracting privacy policy paragraphs (termed segments in this paper) and individual sentences that relate to expert-identified categories of policy contents, using methods in supervised learning. In particular, we show that relevant segments and sentences can be classified with average micro-F1 scores of 0.78 and 0.66 respectively, improving over prior work. We discuss how the techniques introduced in this paper have been used to automatically annotate the text of about 7,000 privacy policies. Our discussion highlights opportunities as well as limitations associated with our classification approach.