Quantifying the Effect of In-Domain Distributed Word Representations : A Study of Privacy Policies
Quantifying the Effect of In-Domain Distributed Word Representations : A Study of Privacy Policies
复制标题
量化域内分布式词表示的效果:隐私政策研究
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
N. Sadeh
中科院分区:
文献类型:
--
作者:
Vinayshekhar Bannihatti Kumar;Abhilasha Ravichander;Peter Story;N. Sadeh
Privacy policies are documents that describe what data is collected by a website or an app and how that data is handled. Privacy policies are often long and difficult to understand. Recently people have started to turn to Natural Language Processing (NLP) to automatically extract statements from the text of these policies. This article reports on a study to evaluate the benefits of using word embeddings in this endeavor. Specifically, we use 150,000 privacy policies to build word vectors in an unsupervised manner. This includes evaluating the benefits of privacy specific word embeddings. Evaluation is conducted on the OPP-115 corpus of privacy policy annotations. By building privacy-specific embeddings we hope to accelerate research at the intersection of privacy policies and language technologies.