Quantifying the Effect of In-Domain Distributed Word Representations : A Study of Privacy Policies

Quantifying the Effect of In-Domain Distributed Word Representations : A Study of Privacy Policies
复制标题

量化域内分布式词表示的效果:隐私政策研究

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
N. Sadeh
N. Sadeh
中科院分区:
--
文献类型:
--
作者:
Vinayshekhar Bannihatti Kumar;Abhilasha Ravichander;Peter Story;N. Sadeh

文献摘要

被引文献

相似文献

隐私政策是描述网站或应用程序收集哪些数据以及如何处理这些数据的文档。隐私政策通常很长,很难理解。最近,人们开始转向自然语言处理(NLP)来自动从这些策略的文本中提取语句。本文报道了一项研究,以评估在这奋进使用词嵌入的好处。具体来说,我们使用150,000个隐私策略以无监督的方式构建词向量。这包括评估隐私特定词嵌入的好处。评估是在隐私政策注释的CN-115语料库上进行的。通过构建特定于隐私的嵌入,我们希望加速隐私政策和语言技术交叉领域的研究。
Privacy policies are documents that describe what data is collected by a website or an app and how that data is handled. Privacy policies are often long and difficult to understand. Recently people have started to turn to Natural Language Processing (NLP) to automatically extract statements from the text of these policies. This article reports on a study to evaluate the benefits of using word embeddings in this endeavor. Specifically, we use 150,000 privacy policies to build word vectors in an unsupervised manner. This includes evaluating the benefits of privacy specific word embeddings. Evaluation is conducted on the OPP-115 corpus of privacy policy annotations. By building privacy-specific embeddings we hope to accelerate research at the intersection of privacy policies and language technologies.