Preserving Privacy in Medical Data Sets
Preserving Privacy in Medical Data Sets
批准号:
6733529
负责人:
STAAL A VINTERBO
金额:
$40.7万
依托单位国家:
美国
项目类别:
财政年份:
2002
资助国家:
美国
项目状态:
已结题
起止时间:
2002-02-01 至 2005-07-31
关键词:
Internetbehavioral /social science research tagcomputer program /softwarecomputer simulationcomputer system design /evaluationconfidentialitydata managementdecision makinghealth care facility information systemhealth care policyhuman datahuman rightsinformation disseminationinformation retrievalmathematical modelmedical recordsmodel design /developmentpatient oriented researchstatistics /biometry
中文摘要
隐私是一项基本权利,需要得到保护。对于卫生保健相关信息,有披露的规定。这些规定的动机是公众对可能导致歧视的违反保密规定的担忧。电子病历技术、互联网和基因革命的最新进展,以及媒体对侵犯隐私的报道,引起了人们对这一话题越来越多的兴趣。人们普遍认为,随着联网计算机的使用,敏感信息更容易获得。由于完全不披露是不现实的,目前的法规要求向某一方提供“最少量”的信息。缺乏对特定类型应用的“最低”构成要素和“有用性指数”的透彻研究。对未识别或匿名的数据库中侵犯隐私的可能性也缺乏准确的量化。这些指标的定义和量化对于决策非常重要。正如我们所证明的,未识别的数据集仍然可以用于推断,因此可能会泄露敏感信息。使用机器学习方法来验证去识别数据集中的剩余函数依赖关系导致更好地理解可能的推论。基于逻辑学、统计学、数据库理论和机器学习方法的匿名技术可以帮助保护隐私。我们将从理论和实践的角度正式定义和研究数据库中的匿名性。我们将开发和实施算法来匿名化数据集,这将符合匿名性和已披露数据集的“有用性”之间的平衡。我们还将开发和实施算法,以验证给定数据集的匿名性,并指示隐私攻击风险最高的记录类型。我们将通过WWW向研究人员免费提供我们的方法和有文档记录的工具。
英文摘要
Privacy is a fundamental right and needs to be protected. For health care related d information, there are regulations for disclosure. These regulations were motivated by the public's concern of breaches of confidentiality that might result in discrimination. The recent progress in electronic medical record technology, the Internet, and the genetic revolution, together with media reports on violations of privacy have generated increasing interest in this topic. A common belief is that sensitive information is more easily available with the use of networked computers. Since total lack of disclosure is not realistic, current regulations require that the "minimal amount" of information be given to a certain party. A thorough study on what constitutes "minimal" for particular types of applications and a "usefulness index" is lacking. An exact quantification of the potential for privacy breach in de-identified or anonymized databases is also lacking. Definition and quantification of these indices is important for decision-making. As we demonstrate, de-identified data sets can still be used for inference and therefore may disclose sensitive information. The use of machine learning methods to verify the remaining functional dependencies in a de- identified data set leads to better understanding of the possible inferences. Anonymization techniques based on logic, statistics, database theory, and machine learning methods can help in the protection of privacy. We will formally define and study anonymity in databases, from a theoretical and a practical standpoint. We will develop and implement algorithms to anonymize data sets that will be in accordance with the balance of anonymity and "usefulness" of the disclosed data sets. We will also develop and implement algorithms to verify the anonymity of a given data set and indicate the type of records that are at highest risk for a privacy attack. We will make our methods and documented tools freely available to researchers via the WWW.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Preserving Privacy in Medical Data Sets
-
批准号:7285699
-
项目类别:
-
资助金额:$29.74万
-
财政年份:2006
-
负责人:STAAL A VINTERBO
-
依托单位:
Preserving Privacy in Medical Data Sets
-
批准号:7143725
-
项目类别:
-
资助金额:$35.0万
-
财政年份:2006
-
负责人:STAAL A VINTERBO
-
依托单位:
Preserving Privacy in Medical Data Sets
-
批准号:6421732
-
项目类别:
-
资助金额:$38.44万
-
财政年份:2002
-
负责人:STAAL A VINTERBO
-
依托单位:
Preserving Privacy in Medical Data Sets
-
批准号:6620783
-
项目类别:
-
资助金额:$38.08万
-
财政年份:2002
-
负责人:STAAL A VINTERBO
-
依托单位: