Towards Practical Federated Analytics and Multi-Target Privacy Enhancing Technologies (PETs)
Towards Practical Federated Analytics and Multi-Target Privacy Enhancing Technologies (PETs)
批准号:
2598750
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
隐私增强技术(或称PETS)是允许隐私数据收集、隐私数据分析和隐私保护机器学习的技术。常见的宠物包括差分隐私、安全多方计算和同态加密。虽然这些工具可以提供很多东西,但这些技术在实践中的采用率很低,而且它们不容易应用于参与者众多、处理能力有限、通信带宽较小的现代场景。这些问题对行业的广泛采用和大规模部署构成了障碍,这将使隐私保护数据科学成为可能。因此,迫切需要在宠物方面取得突破,以兑现保护隐私的机器学习和私人数据分析的承诺。该领域的最新进展是基于将PETS与联邦学习(FL)概念相结合的想法。FL的核心思想很简单:使多方能够在不共享任何本地(私有)数据的情况下共同训练机器学习模型。最近,出现了更广泛的联合分析(FA)概念。FA的目的是采用FL的核心思想,由客户对其数据执行本地计算,并仅将聚集的结果提供给中央服务器。与FL不同,它的重点不是学习,而是收集客户的直接分析。该博士学位的目标是在FL和FA环境中大幅扩展隐私保护联邦计算的可能性,并为使端到端隐私保护数据科学管道成为可行的解决方案做出贡献。该项目的主要重点将是开发新的FA技术。由于FA是一个相对较新的领域,从计算特定分析(如分位数和范围查询)到更一般地处理任意聚合计算,有很多开放方向。一个关键方向将是创造技术,允许我们将FA方法与FL相结合,形成端到端的私人学习系统和私人数据分析管道。这是使这些技术在工业上可行并最终得到广泛采用的一个重要方向。研究方向包括:1.随时间收集分析:分析的一个重要应用是收集一段时间内用户组的数据并分析时间趋势。如果不小心处理,收集同一用户的数据可能会降低隐私。最近的结果允许在动态数据库和时间序列数据中进行私人时间统计。将这一点推广到FA案件,将是使这些数据收集技术对行业实用的另一个重要步骤。安全聚合和分布式DP:FL和FA是笼统的术语,它们可以在各种隐私模型中实例化。例如,安全聚合,其中使用加密方法组合来自所有用户的输入以计算准确答案;以及分布式差分保密,其中用户在其输入中添加DP噪声的“份额”。一个研究方向是研究这些隐私模型及其局限性,并提出新的技术来处理它们。使用PETS支持FL压缩方法:有效的FL方法取决于确保参与的客户端的通信开销很小,因为这一成本可能是FL系统的主要瓶颈。这些方法通常将标准FL技术与模型更新的量化相结合来压缩数据。然而,不太清楚如何将上述隐私模型扩展到支持压缩或量化通信。因此,有必要设计与这些隐私模型兼容的压缩方法。与EPSRC研究主题保持一致:ICT网络和分布式系统;人工智能(值得信赖的自主系统);数字经济(信任、身份、隐私和安全);全球不确定性(网络安全)
英文摘要
Privacy Enhancing Technologies (or PETs) are techniques that allow for private data collection, private data analytics and privacy-preserving machine learning. Popular PETs include Differential Privacy, Secure Multiparty Computation and Homomorphic Encryption. While these tools have much to offer, adoption of these techniques in practice is low and they do not easily apply to modern scenarios where there are many participants, often with limited processing power and small communication bandwidth. These issues pose a barrier to wide-spread industry adoption and large-scale deployments which would make privacy-preserving data science possible. There is therefore an urgent demand for breakthroughs in PETs to deliver on the promise of privacy-preserving machine learning and private data analysis. Recent advances in the area are based on the idea of combining PETs with the concept of Federated Learning (FL). The core idea of FL is simple: To enable multiple parties to jointly train a machine learning model without sharing any local (private) data. More recently, the broader notion of Federated Analytics (FA) has emerged. The aim of FA is to take the core ideas of FL, with client's performing local computations over their data and making only the aggregated results available to a central server. Unlike FL, the focus is less on learning and more on collecting direct analytics about clients. The objectives for this PhD are to substantially extend what is possible with privacy-preserving federated computation in both FL and FA settings and to contribute towards making end-to-end privacy-preserving data science pipelines a viable solution. The main focus of this project will be to develop new FA techniques. As FA is a relatively new field there are plenty of open directions from computing specific analytics such as quantiles and range queries to more generally addressing arbitrary aggregate computations. A key direction will be creating techniques that allow us to combine FA methods with FL to form end-to-end private learning systems and private data analytics pipelines. This is a vital direction for making these technologies viable for industry use and for eventual widespread adoption. Research directions include:1. Collecting analytics over time: An important application of analytics is collecting data on a user-group over time and analysing temporal trends. Collecting data on the same user can degrade privacy if not handled carefully. Recent results have allowed for private temporal statistics in dynamic databases and time series data. Extending this to the FA case would be another important step in making these data collection techniques practical for industry.2. Secure Aggregation and Distributed DP: FL and FA are blanket terms, they can be instantiated within a variety of privacy models. Such as secure aggregation, where cryptographic methods are used to combine the inputs from all users to compute the exact answer, and distributed differential privacy, where users add a "share" of DP noise to their input. One research direction would be to study these privacy models and their limitations and propose new techniques to deal with them.3. Supporting FL compression methods with PETs: Effective FL methods hinge on ensuring the communication overhead of clients participating is small, since this cost can be a primary bottleneck for FL systems. Methods often combine the standard FL techniques with quantization of model updates to compress data. However, it is less clear how the aforementioned privacy models can be extended to support compressed or quantized communications. Hence there is a need for the design of compression methods that are compatible with these privacy models.Alignment with EPSRC research themes: ICT networks and distributed systems; Artificial Intelligence (Trustworthy Autonomous Systems); Digital Economy (Trust, identity, privacy and security); Global Uncertainties (Cybersecurity)
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金