Achieving usable and privacy-assured similarity search over outsourced cloud data

Achieving usable and privacy-assured similarity search over outsourced cloud data
复制标题

DOI:
10.1109/infcom.2012.6195784
复制
发表时间:
2012-03
期刊:
2012 Proceedings IEEE INFOCOM
影响因子:
--
通讯作者:
Cong Wang-;K. Ren;Shucheng Yu;Karthik Mahendra Raje Urs
Cong Wang-;K. Ren;Shucheng Yu;Karthik Mahendra Raje Urs
中科院分区:
其他
文献类型:
--
作者:
Cong Wang-;K. Ren;Shucheng Yu;Karthik Mahendra Raje Urs

文献摘要

被引文献

相似文献

随着个人和企业产生的需要存储和利用的数据迅速增加,数据所有者因其极大的灵活性和经济性而将其本地复杂的数据管理系统外包到云中。然而,由于敏感的云数据可能需要在外包前进行加密,这将取代传统的基于明文关键字搜索的数据利用服务,因此如何实现外包云数据的隐私保护利用机制至关重要。考虑到云上大量的按需数据用户和海量的外包数据文件,这一问题尤其具有挑战性,因为同时满足性能、系统可用性和高级用户搜索体验的实际需求极其困难。本文研究了外包云数据的安全高效相似搜索问题。相似度搜索是明文信息检索中广泛使用的一种基本而有力的工具,但在加密数据领域还没有得到很好的探索。我们的机制设计首先利用抑制技术从给定的文档集合中构建存储效率高的相似度关键字集,并以编辑距离作为相似度度量。在此基础上,构建了一个私有Trie遍历搜索索引,并证明了该索引在搜索时间复杂度不变的情况下,正确地实现了定义的相似搜索功能。在严格的安全处理下,我们形式化地证明了该机制的隐私保护保证。为了展示我们机制的通用性并进一步丰富应用范围,我们还展示了我们的新结构自然支持模糊搜索,这是一个以前研究过的概念,旨在容忍用户搜索输入中的打字错误和表示不一致。在Amazon云平台上使用真实数据集进行的大量实验进一步验证了该机制的有效性和实用性。
As the data produced by individuals and enterprises that need to be stored and utilized are rapidly increasing, data owners are motivated to outsource their local complex data management systems into the cloud for its great flexibility and economic savings. However, as sensitive cloud data may have to be encrypted before outsourcing, which obsoletes the traditional data utilization service based on plaintext keyword search, how to enable privacy-assured utilization mechanisms for outsourced cloud data is thus of paramount importance. Considering the large number of on-demand data users and huge amount of outsourced data files in cloud, the problem is particularly challenging, as it is extremely difficult to meet also the practical requirements of performance, system usability, and high-level user searching experiences. In this paper, we investigate the problem of secure and efficient similarity search over outsourced cloud data. Similarity search is a fundamental and powerful tool widely used in plaintext information retrieval, but has not been quite explored in the encrypted data domain. Our mechanism design first exploits a suppressing technique to build storage-efficient similarity keyword set from a given document collection, with edit distance as the similarity metric. Based on that, we then build a private trie-traverse searching index, and show it correctly achieves the defined similarity search functionality with constant search time complexity. We formally prove the privacy-preserving guarantee of the proposed mechanism under rigorous security treatment. To demonstrate the generality of our mechanism and further enrich the application spectrum, we also show our new construction naturally supports fuzzy search, a previously studied notion aiming only to tolerate typos and representation inconsistencies in the user searching input. The extensive experiments on Amazon cloud platform with real data set further demonstrate the validity and practicality of the proposed mechanism.