kb-anonymity: a model for anonymized behaviour-preserving test and debugging data

kb-anonymity: a model for anonymized behaviour-preserving test and debugging data
复制标题

DOI:
10.1145/1993498.1993551
复制
发表时间:
2011-06
期刊:
--
影响因子:
--
通讯作者:
Aditya Budi;D. Lo;Lingxiao Jiang;Lucia
Aditya Budi;D. Lo;Lingxiao Jiang;Lucia
中科院分区:
其他
文献类型:
--
作者:
Aditya Budi;D. Lo;Lingxiao Jiang;Lucia

文献摘要

被引文献

相似文献

生成可以执行程序中所有可能的程序状态的测试用例通常非常昂贵并且实际上不可行。对于中型或大型工业系统尤其如此。在实践中,系统的工业客户通常在系统构建之前或系统的先前版本部署之后收集一组输入数据。此类数据非常有价值,因为它们代表了客户日常业务中重要的操作,并且可用于广泛测试系统。然而,此类数据通常携带敏感信息,无法发布给第三方开发公司。例如,医疗保健提供者可能拥有一组严格保密且不能被任何第三方使用的患者记录。仅仅简单地屏蔽敏感值可能还不够,因为数据中字段之间的相关性可以揭示被屏蔽的信息。此外,屏蔽数据可能会在系统中表现出不同的行为,并且在测试和调试方面变得不如原始数据有用。为了释放私有数据进行测试和调试,本文提出了kb-匿名模型,该模型将数据挖掘和数据库领域常用的k-匿名模型与程序行为保存的概念相结合。与k-anonymity一样,kb-anonymity会替换原始数据中的部分信息,以确保隐私保护,以便替换后的数据可以发布给第三方开发者。与 k-匿名性不同,kb-匿名性确保替换的数据表现出与原始数据表现出的相同类型的程序行为,以便替换的数据对于测试和调试的目的仍然有用。我们还提供了三种特定配置下的模型的具体版本,并成功地将我们的原型实现应用于三个开源程序,展示了我们原型的实用性和可扩展性。
It is often very expensive and practically infeasible to generate test cases that can exercise all possible program states in a program. This is especially true for a medium or large industrial system. In practice, industrial clients of the system often have a set of input data collected either before the system is built or after the deployment of a previous version of the system. Such data are highly valuable as they represent the operations that matter in a client's daily business and may be used to extensively test the system. However, such data often carries sensitive information and cannot be released to third-party development houses. For example, a healthcare provider may have a set of patient records that are strictly confidential and cannot be used by any third party. Simply masking sensitive values alone may not be sufficient, as the correlation among fields in the data can reveal the masked information. Also, masked data may exhibit different behavior in the system and become less useful than the original data for testing and debugging. For the purpose of releasing private data for testing and debugging, this paper proposes the kb-anonymity model, which combines the k-anonymity model commonly used in the data mining and database areas with the concept of program behavior preservation. Like k-anonymity, kb-anonymity replaces some information in the original data to ensure privacy preservation so that the replaced data can be released to third-party developers. Unlike k-anonymity, kb-anonymity ensures that the replaced data exhibits the same kind of program behavior exhibited by the original data so that the replaced data may still be useful for the purposes of testing and debugging. We also provide a concrete version of the model under three particular configurations and have successfully applied our prototype implementation to three open source programs, demonstrating the utility and scalability of our prototype.