Financial crime prevention using privacy preserving Federated Learning
Financial crime prevention using privacy preserving Federated Learning
批准号:
10048524
负责人:
金额:
$1.27万
依托单位:
依托单位国家:
英国
项目类别:
CR&D Bilateral
财政年份:
2022
资助国家:
英国
项目状态:
已结题
起止时间:
2022 至 --
中文摘要
金融犯罪和金融欺诈每年给银行和相关公司造成数十亿美元的损失。建立一个欺诈检测机器学习模型来防止这种欺诈交易是可能的,但由于缺乏相关的训练数据,其使用受到限制。金融公司拥有大量的交易数据,但由于隐私和安全问题,这些数据无法使用。联邦学习(FL)是一种机器学习(ML)技术,可以从这些金融公司的私人数据中训练算法,而不会损害用户/客户的隐私。与传统的ML方法相比,FL自然提供了隐私优势。然而,对于高度敏感的数据,FL周期的每一步都需要隐私,以确保私人数据不会通过训练的ML模型泄露。在本文中,我们提出了一种隐私保护的联邦学习(FL)解决方案,以解决高度敏感的银行交易数据中的欺诈检测问题。我们的联邦学习隐私技术堆栈由 ** 启用SSL的服务器-客户端 ** 连接和 ** SecAgg +**(安全聚合+)策略组成,用于屏蔽本地客户端模型。我们的安全聚合策略是稳健的客户端退出。我们通过在训练梯度中引入噪声,进一步加强了我们的隐私堆栈。根据所需任务的隐私-效用权衡,可以可选地添加或删除差异隐私。对于我们的欺诈/异常检测任务,我们提出了一个 ** FT-transformer **,它与FL框架兼容,并与XGBoost和随机森林等表格模型相媲美。通过结合最先进的隐私算法和深度学习架构,我们创建了一个独特的,可扩展的,可定制的,机器学习任务无关的,ML框架无关的和隐私保护的联邦学习解决方案。
英文摘要
Financial crime and financial fraud cost billions of dollars to banks and related firms every year. It is possible to build a fraud detection machine learning model to prevent such fraudulent transactions, but due to a lack of relevant training data, its use is limited. There is a wealth of transactional data that financial firms possess, but due to privacy and security concerns, the data cannot be utilized. Federated learning(FL) is a machine learning (ML) technique that can train an algorithm from the private data of these financial firms without compromising user/client's privacy. FL naturally offers privacy advantages compared to the traditional ML approaches. However for highly sensitive data, privacy at each step of FL cycle is required to ensure that private data is not leaked through trained ML models.In this paper, we proposed a privacy-preserving federated learning (FL) solution to tackle fraud detection in highly sensitive bank transactional data. Our federated learning privacy technology stack consists of **SSL-enabled server-client** connection and a **SecAgg+** (secure aggregation plus) strategy to mask local client models. Our secure aggregation strategy is robust client dropouts. We further strengthened our privacy stack with **differential privacy** by introducing noise in the training gradients. Differential privacy can be optionally added or removed depending on the privacy-utility trade-off of the required task. For our fraud/anomaly detection task, we have proposed a **FT-transformer** which is compatible with FL framework and performs at par with tabular models such as XGBoost and random forest. By combining state-of-the-art privacy algorithms and deep learning architecture, we created a unique, scalable, customisable, machine learning task-agnostic, ML framework agnostic and privacy-preserving federated learning solution.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金