You are your Metadata: Identification and Obfuscation of Social Media Users using Metadata Information

You are your Metadata: Identification and Obfuscation of Social Media Users using Metadata Information
复制标题

DOI:
10.1609/icwsm.v12i1.15010
复制
发表时间:
2018-03
期刊:
--
影响因子:
--
通讯作者:
Beatrice Perez;Mirco Musolesi;G. Stringhini
Beatrice Perez;Mirco Musolesi;G. Stringhini
中科院分区:
其他
文献类型:
--
作者:
Beatrice Perez;Mirco Musolesi;G. Stringhini

文献摘要

被引文献

相似文献

元数据与我们在数字世界中的日常互动和交流中产生的大多数信息相关联。然而,令人惊讶的是,元数据通常仍然被归类为非敏感。事实上,过去,研究人员和从业者主要关注的是从消息内容中识别用户的问题。在本文中,我们使用Twitter作为案例研究,量化元数据和用户身份之间的关联的唯一性,并了解潜在的混淆策略的有效性。更具体地说,我们分析元数据中的原子字段,并系统地将它们联合收割机组合起来,以便使用不同的机器学习算法将新推文分类为属于一个帐户。我们证明,通过应用监督学习算法,我们能够以约96.7%的准确率识别10,000人组中的任何用户。此外,如果我们扩大搜索范围并考虑10个最有可能的候选人,我们将模型的准确率提高到99.22%。我们还发现,数据混淆对于这类数据来说是困难和无效的:即使在干扰了60%的训练数据之后,仍然可以以高于95%的准确率对用户进行分类。这些结果有很强的影响,在元数据混淆策略的设计,例如数据集的发布,不仅为Twitter,但更普遍的是,大多数社交媒体平台。
Metadata are associated to most of the information we produce in our daily interactions and communication in the digital world. Yet, surprisingly, metadata are often still categorized as non-sensitive. Indeed, in the past, researchers and practitioners have mainly focused on the problem of the identification of a user from the content of a message. In this paper, we use Twitter as a case study to quantify the uniqueness of the association between metadata and user identity and to understand the effectiveness of potential obfuscation strategies. More specifically, we analyze atomic fields in the metadata and systematically combine them in an effort to classify new tweets as belonging to an account using different machine learning algorithms of increasing complexity. We demonstrate that, through the application of a supervised learning algorithm, we are able to identify any user in a group of 10,000 with approximately 96.7% accuracy. Moreover, if we broaden the scope of our search and consider the 10 most likely candidates we increase the accuracy of the model to 99.22%. We also found that data obfuscation is hard and ineffective for this type of data: even after perturbing 60% of the training data, it is still possible to classify users with an accuracy higher than 95%. These results have strong implications in terms of the design of metadata obfuscation strategies, for example for data set release, not only for Twitter, but, more generally, for most social media platforms.