Combatting The Challenges of Local Privacy for Distributional Semantics with Compression

Combatting The Challenges of Local Privacy for Distributional Semantics with Compression
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Alexandra Schofield
Alexandra Schofield
中科院分区:
其他
文献类型:
--
作者:
Alexandra Schofield

文献摘要

相似文献

在词袋特征中加入局部私有噪声的传统方法掩盖了文本数据中的真实信号,从而消除了分布式语义模型所依赖的稀疏性和非否定性。我们认为有限精度局部隐私的表述是一个更合适的词袋特征框架,它保证了小于用户指定的最大距离的文档之间的隐私。为了减少必须添加随机噪声的特征数量,我们还在添加噪声之前压缩单词特征,然后在模型推理之前解压缩这些特征。我们测试了随机的聚合方法以及由单词的分布属性通知的方法。将LDA和LSA应用于合成数据和真实数据,我们发现这些方法产生的分布模型更接近原始数据。
Traditional methods for adding locally private noise to bag-of-words features overwhelm the true signal in the text data, removing the properties of sparsity and non-negativity often relied upon by distributional semantic models. We argue the formulation of limited-precision local privacy, which guarantees privacy between documents of less than a user-specified maximum distance, is a more appropriate framework for bag-of-words features. To reduce the number of features to which we must add random noise, we also compress word features before adding noise, then decompress those features before model inference. We test randomized methods of aggregation as well as methods informed by distributional properties of words. Applying LDA and LSA to synthetic and real data, we show that these approaches produce distributional models closer to those in the original data.