Enabling Health Data Sharing with Fine-Grained Privacy

Enabling Health Data Sharing with Fine-Grained Privacy
复制标题

DOI:
10.1145/3583780.3614864
复制
发表时间:
2023-10
期刊:
Proceedings of the 32nd ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Luca Bonomi;Sepand Gousheh;Liyue Fan
Luca Bonomi;Sepand Gousheh;Liyue Fan
中科院分区:
其他
文献类型:
--
作者:
Luca Bonomi;Sepand Gousheh;Liyue Fan

文献摘要

相似文献

共享健康数据对于推进医学研究并将知识转化为临床实践至关重要。同时,保护数据贡献者的隐私至关重要。为此,已经提出了几种隐私方法来保护数据共享中的单个数据贡献者,包括数据匿名和数据合成技术。这些方法在在数据集级别提供隐私保护方面显示出令人鼓舞的结果。在这项工作中,我们研究了在健康数据共享中实现细粒度隐私方面面临的隐私挑战。我们的工作是由最近的研究结果激励的,在这些发现中,患者和医疗保健提供者可能具有需要解决的不同隐私偏好和政策。具体而言,我们提出了一种新颖有效的隐私解决方案,该解决方案使数据策展人(例如医疗保健提供者)能够保护敏感的数据元素,同时保留数据有用性。我们的解决方案以随机技术为基础,为敏感元素提供严格的隐私保护,并利用图形模型来减轻因依赖元素而引起的隐私泄漏。为了增强共享数据的实用性,我们的随机机制结合了域知识以保持语义相似性并采用块结构化设计以最大程度地减少效用损失。使用现实世界中健康数据的评估证明了我们的方法的有效性以及共享数据对健康应用程序的实用性。
Sharing health data is vital in advancing medical research and transforming knowledge into clinical practice. Meanwhile, protecting the privacy of data contributors is of paramount importance. To that end, several privacy approaches have been proposed to protect individual data contributors in data sharing, including data anonymization and data synthesis techniques. These approaches have shown promising results in providing privacy protection at the dataset level. In this work, we study the privacy challenges in enabling fine-grained privacy in health data sharing. Our work is motivated by recent research findings, in which patients and healthcare providers may have different privacy preferences and policies that need to be addressed. Specifically, we propose a novel and effective privacy solution that enables data curators (e.g., healthcare providers) to protect sensitive data elements while preserving data usefulness. Our solution builds on randomized techniques to provide rigorous privacy protection for sensitive elements and leverages graphical models to mitigate privacy leakage due to dependent elements. To enhance the usefulness of the shared data, our randomized mechanism incorporates domain knowledge to preserve semantic similarity and adopts a block-structured design to minimize utility loss. Evaluations with real-world health data demonstrate the effectiveness of our approach and the usefulness of the shared data for health applications.