课题基金 / 基金详情

Balancing Disclosure Risk with Inferential Power: Software for Intervalized Data

Balancing Disclosure Risk with Inferential Power: Software for Intervalized Data
平衡披露风险与推理能力:间隔数据软件
批准号:
8517848
负责人:
SCOTT D FERSON
金额:
$23.61万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-08-01 至 2014-05-31

项目摘要

项目成果

SCOTT D FERSON的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):在医疗保健提供和公共卫生调查期间收集的患者数据拥有大量可用于生物医学和流行病学研究的信息。然而,由于大多数个人健康记录的隐私性质,对这些数据的访问通常是有限的。在这些数据可用于改善公共卫生之前,需要采取方法平衡研究数据的信息量与最大限度地减少披露风险所需的信息损失。目前的方法主要集中在保护隐私上,但仅仅集中在保护隐私上是不够的。在统计披露控制技术中,信息真实性没有得到很好的保护,从而可能发布不可靠的结果。在基于泛化的匿名化方法中,存在由于属性泛化而导致的信息丢失,并且现有技术没有提供足够的控制来维护数据效用。目前需要的是既保护数据中所代表的个人隐私,又保护研究人员所研究的关系完整性的方法。问题是,在保护个人隐私和保护数据集的信息性之间存在固有的权衡。保护个人隐私总是会导致信息的丢失,而正是数据集所包含的信息影响了统计测试的效力。然而,对于给定的匿名化策略,通常有多种方法来掩盖满足所提供的披露风险标准的数据。可以利用这一点来选择既能最好地保存统计信息又能满足所提供的披露风险标准的解决方案。该项目将开发第一个集成软件系统,为敏感医疗数据发布的所有三个阶段所面临的问题提供解决方案:1。通过使数据间隔化/一般化以满足当前可用的匿名化策略来使数据集匿名化,2.在匿名化程序中提供足够的控制,以满足对数据的统计有用性的约束,以及3.计算匿名数据区间的统计检验。这项工作面临两大挑战。首先,根据现有的研究结果,将我们提出的新控制过程集成到匿名化过程中预计在计算上是困难的。我们将通过开发高效且实用的贪婪算法、近似算法或适用于实际情况的算法(如果不是一般情况)来克服这一挑战。这项工作面临的另一个主要挑战是,已知区间数据集的统计计算在计算上是困难的,并且这些计算对于匿名化过程中的控制过程以及后续的统计计算和测试都是必要的。我们将克服这一挑战,有效的算法,利用数据集的结构间隔隐私。该软件将在各种大小和结构的医疗数据集上进行测试,以证明该方法的可行性,并表征数据集大小的算法的可扩展性。
英文摘要
DESCRIPTION (provided by applicant): Patient data collected during health care delivery and public health surveys possess a great deal of information that could be used in biomedical and epidemiological research. Access to these data, however, is usually limited because of the private nature of most personal health records. Methods of balancing the informativeness of data for research with the information loss required to minimize disclosure risk are needed before these data can be used to improve public health. Current methods are primarily focused on protecting privacy, but focusing on protecting privacy alone is inadequate. In statistical disclosure control techniques, information truthfulness is not well preserved so that unreliable results may be released. In generalization-based anonymization approaches, there is information loss due to attribute generalization and existing techniques do not provide sufficient control for maintaining data utility. What are currently needed are methods that protect both the privacy of individuals represented in the data as well as the integrity of relationships studied by researchers. The problem is that there is an inherent tradeoff between protecting the privacy of individuals and protecting the informativeness of the data set. Protecting the privacy of individuals always results in a loss of information and it is the information contained by the data set that affects the power of a statistical test. For a given anonymization strategy, however, there are often multiple ways of masking the data that meet the disclosure risk criteria provided. This can be taken advantage of to choose the solution that best preserves statistical information while meeting the disclosure risk criteria provided. This project will develop the first integrated software system that provides solutions for problems faced in all three stages in the release of sensitive health care data: 1. anonymize a data set by intervalizing/generalizing data to satisfy currently available anonymization strategies, 2. provide sufficient controls within anonymization procedures to satisfy constraints on statistical usefulness of the data, and 3. compute statistical tests for the anonymized data intervals. There are two main challenges facing this effort. The first is that, based on existing research results, integrating our proposed new control processes into anonymization procedures is expected to be computationally difficult. We will overcome this challenge by developing efficient and practically useful greedy algorithms, approximation algorithms, or algorithms working for realistic situations (if not for general cases). The other primary challenge facing this effort is the fact that statistical calculations with interval data sets are known to be computationally difficult, and these calculations are necessary both for control processes within anonymization procedures and for subsequent statistical computation and tests. We will overcome this challenge with efficient algorithms that exploit the structure present in data sets intervalized for privacy. The software will be tested on medical data sets of various sizes and structures to demonstrate the feasibility of the approach and to characterize the scalability of the algorithms with data set size.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Data Anonymization that Leads to the Most Accurate Estimates of Statistical Characteristics: Fuzzy-Motivated Approach.
数据匿名化可实现最准确的统计特征估计:模糊驱动方法。
DOI: 10.1109/ifsa-nafips.2013.6608471
发表时间: 2013
期刊: Proceedings. IFSA World Congress
影响因子: --
作者: [Xiang,G, Ferson,S, Ginzburg,L, Longpré,L, Mayorga,E, Kosheleva,O]
通讯作者: Kosheleva,O
Balancing Disclosure Risk with Inferential Power: Software for Intervalized Data
  • 批准号:
    8251091
  • 项目类别:
  • 资助金额:
    $24.47万
  • 财政年份:
    2012
  • 负责人:
    SCOTT D FERSON
  • 依托单位:
Compensating for Uncertainty Biases in Health Risk Judgments
  • 批准号:
    7926647
  • 项目类别:
  • 资助金额:
    $162.78万
  • 财政年份:
    2010
  • 负责人:
    SCOTT D FERSON
  • 依托单位:
Safe environmental concentrations under uncertainty
  • 批准号:
    6337570
  • 项目类别:
  • 资助金额:
    $9.99万
  • 财政年份:
    2001
  • 负责人:
    SCOTT D FERSON
  • 依托单位:
Safe environment concentrations under uncertainty
  • 批准号:
    6788050
  • 项目类别:
  • 资助金额:
    $34.69万
  • 财政年份:
    2000
  • 负责人:
    SCOTT D FERSON
  • 依托单位:
海外基金