Privacy Aware Temporal Profiling of Emails in Distributed Setup

Privacy Aware Temporal Profiling of Emails in Distributed Setup
复制标题

分布式设置中电子邮件的隐私意识时间分析

DOI:
10.1145/3132847.3132970
复制
发表时间:
2017
期刊:
Proceedings of the 2017 ACM on Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
S. Lodha
S. Lodha
中科院分区:
--
文献类型:
--
作者:
Sutapa Mondal;Manish Shukla;S. Lodha

文献摘要

被引文献

相似文献

企业电子邮件有望成为知识发现的丰富来源。这是由于通信的直接性质、对不同媒体类型的支持、实体的积极参与以及按时间顺序排列的消息。此外,由于企业电子邮件的正式性质,它比外部电子邮件更值得信任。该数据源尚未得到充分利用。事实上,现有的电子邮件分析工作主要集中在专业知识识别和检索上。即使在这些研究中,研究人员也做出了一些限制性的假设。例如,在许多公式中,底层系统假设一个集中的数据存储库,并且通信网络是完整的。在挖掘和汇总结果时,它们不会考虑电子邮件中的个人偏见。此外,电子邮件包含相当数量的个人和组织敏感信息。现有的电子邮件分析工作都没有提出任何关于缓解个人和组织隐私担忧的建议。在这篇文章中,我们提出了一个系统,用于建立个人的感知知识概况“她知道什么?”),趋势概况“她的专业知识在哪个方向上增长了多少?”),以及团队概况“她所有的队友都知道什么?”)。拟议的系统在分布式网络中运行,并对驻留在时变本地电子邮件数据库中的电子邮件执行分析,而不需要事先对环境进行假设。它还通过从感知到的对等点的简档和它们的共同兴趣来推断它们的简档,从而照顾到部分通信网络中缺失的节点。我们开发了一种两遍聚合算法,用于组合来自各个节点的结果并得出有用的见解。为了进一步提高聚集算法的输出,使用了基于图的算法来计算扩散(REACH)和流行度(RECALL)。结果表明,两遍聚合步骤在大多数情况下是足够的,并且电子邮件内容和基于图的方法的混合在分布式设置中工作得很好。
The enterprise email promises to be a rich source for knowledge discovery. This is made possible due to the direct nature of communication, support for diverse media types, active participation of entities and presence of chronological ordering of messages. Also, the enterprise emails are more trustworthy than external emails due to their formal nature. This data source has not been fully tapped. In fact, the existing work on profiling of emails focuses primarily on expertise identification and retrieval. Even in these studies, the researchers have made some restrictive assumptions. For instance, in many of the formulations, the underlying system assumes a centralized data repository, and the communication network is complete. They do not account for individual biases in an email while mining and aggregating results. Furthermore, email holds fair amount of personal and organizational sensitive information. None of the existing work on email profiling suggests anything on alleviating the individual and organizational privacy concerns. In this paper, we propose a system for building an individual's perceived knowledge profile "What she knows?" ), trends profile "In which direction and how far her expertise has grown?" ), and team profile "What all her teammates know?"). The proposed system operates in a distributed network and performs analysis of emails residing on a time-varying local email database, with no prior assumptions about the environment. It also takes care of missing nodes in a partial communication network, by deducing their profile from perceived profiles of its peers and their common interest. We developed a two-pass aggregation algorithm for combining results from individual nodes and drawing useful insights. A graph based algorithm is used for calculating spread (reach) and popularity (recall) for further improving the output of the aggregation algorithm. The results show that the two pass aggregation step is sufficient in majority of the cases, and a hybrid of email content and graph-based approach works well in a distributed setup.