Theoretical Limits of Data Privacy
Theoretical Limits of Data Privacy
批准号:
2887682
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
该项目属于EPRSC ICT网络和分布式系统研究领域。该项目的重点是使用信息论来研究通信和存储场景中数据隐私的基本限制,它也将更实际地研究我们可以在现实生活系统中接近这些限制的方法。信息论关注的是一个事件或随机变量中信息量的量化,在交流的背景下,这可能是要传达给预期接收者的信息。它可以告诉我们表示一个事件所需的信息量(即比特数),以及该信息可以可靠地通过信道传输的速率。在数据隐私的背景下,本研究的目标将是构建数学定理,指定在各种设置下可以保持数据隐私的条件。例如,这可以回答这样一个问题:“在给定访问某个网络的情况下,数据挖掘者理论上可以收集多少信息?”这里所讨论的网络可以是像用户的社交媒体连接这样简单的东西。数据挖掘导致的数据泄露是当前非常严重的问题。目前,很少有关于在这种事件中可能访问的最大信息量的结果。实现这一结果所采用的方法将涉及对信息不平等和随机编码参数的试验。与遍历数据源相关的证明可能会使用渐近均分性质(AEP),但对于更现实的非遍历数据源,可能需要开发新的方法。该项目的第二个主要目标将是考虑结果的实际意义,并评估是否以及如何在实际系统中接近达到理论极限。在信息论中,有一些基本的界限,我们已经知道很多年了,但在实践中我们仍然没有接近实现。因此,第二个目标与第一个目标截然不同。例如,许多传输边界都是基于无限块长度(即无限长度的码字)的思想推导出来的,从而允许使用AEP。显然,现实生活中的码字不是无限长的,因此实际编码方案所达到的实际速率并不直接遵循理论可实现的界限。回到数据隐私的例子:数据挖掘者按照其当前最佳实用方法收集的信息可能比理论上可能的要少(或多)得多。就应用程序而言,了解在最坏情况下可以获得多少信息,以及实际访问多少信息,以及如何做到这一点,将是非常有用的。
英文摘要
The project falls within the EPRSC ICT Networks and Distributed Systems research area. The focus of the project is the use of information theory to investigate the fundamental limits of data privacy in communications and storage scenario It will also be of interest to look more practically into ways we can come close to approaching those limits in real life systems. Information theory concerns the quantification of the amount of information held in an event or random variable, which in the context of communication could be a message to be conveyed to an intended recipient. It can tell us the amount of information (i.e., number of bits) needed to represent an event, and the rate at which this information can reliably be transmitted across a channel. In the context of data privacy, the goal of this research will be to construct mathematical theorems that specify the conditions under which data privacy can be maintained in various settings. This could for example answer the question: "How much information can a data miner theoretically collect, given access to a certain network?", where the network in question could be something as simple as a user's social media connections. Data breaches as a result of data mining are a very current and serious concern. At present, there are very few results pertaining to the maximum amount of information that could possibly be accessed in such an incident. The methodology employed to achieve such results will involve experimenting with information inequalities and random coding arguments. Proofs related to ergodic data sources are likely to make use of the asymptotic equipartition property (AEP), but new methods may need to be developed for more realistic non-ergodic sources. The second main aim of the project will be to consider the practical meaning of the results, and assess if and how one could come close to achieving the theoretical limits in a real system. In information theory, there exist fundamental bounds that have been known for many years, that we still do not come close to achieving in practice. Thus, this second aim is quite distinct from the first. For example, many transmission bounds are derived with the idea of infinite block lengths (i.e., codewords of infinite length), allowing for the use of the AEP. Clearly, codewords in real life are not infinitely long, so the actual rate achieved by practical coding schemes does not follow directly from the theoretical achievable bound. Returning to the example of data privacy: the information that a data miner can collect following their current best practical methodology may be much less (or more) than what is theoretically possible. In terms of application, it would be useful to know both how much information could be available in a worst-case scenario, as well as how much is realistically accessed, and how this could be done.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金