C‐sanitized: A privacy model for document redaction and sanitization

C‐sanitized: A privacy model for document redaction and sanitization
复制标题

C-sanitized:文档编辑和清理的隐私模型

DOI:
--
复制
发表时间:
2014
期刊:
J. Assoc. Inf. Sci. Technol.
影响因子:
--
通讯作者:
Montserrat Batet
Montserrat Batet
中科院分区:
--
文献类型:
--
作者:
David Sánchez;Montserrat Batet

文献摘要

被引文献

相似文献

每天都会交换和/或发布大量信息。当文档不受控制地提供给不受信任的第三方时,大部分信息的敏感性会造成严重的隐私威胁。在这种情况下,责任组织应采取适当的数据保护措施,特别是在现行数据隐私立法的保护下。为此,通常要求人类专家编辑或清理文档内容。为了减轻这项繁重的任务,本文提出了一种用于文档编辑/清理的隐私模型,与文献中的其他模型相比,该模型具有多种优势。基于数据语义和信息论的完善基础,我们的模型提供了一个框架来开发和实现自动化和固有的语义编辑/清理工具。此外,与临时编辑方法相反,我们的提案提供了先验的隐私保证,可以根据当前的数据隐私立法直观地定义。在几个用例的背景下进行的实证测试说明了我们模型的适用性及其模仿人类消毒剂推理的能力。
Vast amounts of information are daily exchanged and/or released. The sensitive nature of much of this information creates a serious privacy threat when documents are uncontrollably made available to untrusted third parties. In such cases, appropriate data protection measures should be undertaken by the responsible organization, especially under the umbrella of current legislation on data privacy. To do so, human experts are usually requested to redact or sanitize document contents. To relieve this burdensome task, this paper presents a privacy model for document redaction/sanitization, which offers several advantages over other models available in the literature. Based on the well‐established foundations of data semantics and information theory, our model provides a framework to develop and implement automated and inherently semantic redaction/sanitization tools. Moreover, contrary to ad‐hoc redaction methods, our proposal provides a priori privacy guarantees which can be intuitively defined according to current legislations on data privacy. Empirical tests performed within the context of several use cases illustrate the applicability of our model and its ability to mimic the reasoning of human sanitizers.