课题基金 / 基金详情

EthicalML: Injecting Ethical and Legal Constraints into Machine Learning Models

EthicalML: Injecting Ethical and Legal Constraints into Machine Learning Models
EthicalML:将道德和法律约束注入机器学习模型
批准号:
EP/P03442X/1
负责人:
Novi Quadrianto
金额:
$12.83万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
我们对看电影或看小说的选择会受到基于机器学习(ML)的推荐系统的建议的影响。然而,在一些重要的场景中,ML系统是有缺陷的。以下每个场景都涉及我们希望训练ML系统以使其提供服务的情况。然而,在每种情况下,都必须对ML系统的运行施加一个重要的限制。场景1:我们想要一个系统,使提交的工作申请与我们的学术空缺列表相匹配。该制度必须对少数群体一视同仁。场景2:我们需要一个基于活检图像的自动癌症诊断系统。我们还有艾滋病毒检测结果,可以在培训时使用,但不应从新患者身上收集。场景3:我们希望有一个系统,可以帮助我们决定是否批准抵押贷款申请。我们需要了解决策过程,并将其与我们的核对表联系起来,例如申请者是否在过去三个月内透支,是否在选民名册上。情景1要求ML系统在其决定中保持公平,在种族、性别和残疾等方面不歧视;情景2要求ML系统保护个人敏感数据的机密性;情景3要求ML系统通过提供人类可理解的决定来实现透明度。为ML模型配备道德和法律约束(场景1-3)是一个严重的问题;如果没有这一点,ML的未来将处于危险之中。在英国,这一点得到了下议院科学和技术委员会的认可,该委员会建议紧急成立数据伦理委员会(《大数据困境》报告,2016)。此外,自2015年以来,皇家学会启动了一个政策项目,着眼于ML模型及其使用案例的进步所带来的社会、法律和伦理挑战。构建具有公平、保密和透明度约束的ML模型是一个活跃的研究领域,并且有脱节的框架可用于解决每个约束。然而,如何将它们放在一起并不明显。我的长期目标是开发一个具有即插即用约束的ML框架,该框架能够处理所提到的任何约束、它们的组合,以及未来可能规定的新约束。此特权信息在培训时可用,以便更好地训练决策模型并使决策模型具有非歧视性,但在部署时不能供将来的数据访问。对于机密性限制,个人机密数据,如艾滋病毒检测结果,是特权信息。对于公平约束,种族和性别等受保护的特征是特权信息。对于透明度约束,复杂的无法解释但具有高度区分性的特征,如深度学习特征,是特权信息。该项目旨在开发一个ML框架,该框架能够产生准确的预测和对其预测的不确定性估计,同时也符合伦理和法律约束。该方案的主要贡献是:1)一种新的特权学习算法,通过允许在部署时即插即用地使用各种约束,通过被核化,通过优化其超参数,并通过产生预测不确定性的估计,克服了现有方法的局限性;2)可扩展的自动化推理,使得新的特权学习算法很容易适用于任何大规模的学习问题,例如二进制分类、多类分类和回归;以及3)新算法的实例化,用于将公平性、保密性和透明度限制合并到ML模型中。
英文摘要
Our choice as to which movies to watch or novels to read can be influenced by suggestions made by machine learning (ML)-based recommender systems. However, there are some important scenarios where ML systems are deficient. Each of the following scenarios involves a situation where we wish to train an ML system so that it delivers a service. In each case, however, there is an important constraint that must be imposed on the operation of the ML system.Scenario 1: We want a system that will match submitted job applications to our list of academic vacancies. The system has to be non-discriminatory to minority groups. Scenario 2: We need an automated cancer diagnosis system based on biopsy images. We also have HIV test results, which can be used at training time but should not be collected from our new patients.Scenario 3: We wish to have a system that can aid us in deciding whether or not to approve a mortgage application. We need to understand the decision process and relate it to our checklist such as whether or not the applicant has an overdraft in the last three months and is on electoral roll. Scenario 1 asks an ML system to be fair in its decisions by being non-discriminatory with regards to, e.g., race, gender, and disability; scenario 2 requires an ML system to protect confidentiality of personal sensitive data; and scenario 3 demands transparency from an ML system by providing human-understandable decisions. Equipping ML models with ethical and legal constraints, scenarios 1-3, is a serious issue; without this, the future of ML is at risk. In the UK, this is recognized by the House of Commons Science and Technology Committee, which recommended an urgent formation of a Council of Data Ethics ("The Big Data Dilemma" report, 2016). Furthermore, since 2015, the Royal Society has started a policy project that looks at the social, legal, and ethical challenges associated with advancement in ML models and their use cases.Building ML models with fairness, confidentiality, and transparency constraints is an active research area, and disjoint frameworks are available for addressing each constraint. However, how to put them all together is not obvious. My long-term goal is to develop an ML framework with plug-and-play constraints that is able to handle any of the mentioned constraints, their combinations, and also new constraints that might be stipulated in the future.The proposed ML framework relies on instantiating ethical and legal constraints as privileged information. This privileged information is available at training time to better train a decision model and to make a decision model non-discriminatory, but it will not be accessible for future data at deployment time. For confidentiality constraints, personal confidential data such as HIV test results are the privileged information. For fairness constraints, protected characteristics such as race and gender are the privileged information. For transparency constraints, complex un-interpretable but highly discriminative features such as deep learning features are the privileged information.This project aims to develop an ML framework that produces accurate predictions and uncertainty estimates about its predictions while also complying with ethical and legal constraints. The key contributions of this proposal are: 1) a new privileged learning algorithm that overcomes limitations of existing methods by allowing to plug-and-play various constraints at deployment time, by being kernelized, by optimizing its hyperparameters, and by producing estimates of prediction uncertainty, 2) a scalable and automated inference that makes the new privileged learning algorithm easily applicable for any large scale learning problem such as binary classification, multi-class classification, and regression, and 3) an instantiation of the new algorithm for incorporating fairness, confidentiality, and transparency restrictions into ML models.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr.2019.00842
发表时间: 2018-10
期刊: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者: [Novi Quadrianto;V. Sharmanska;Oliver Thomas]
通讯作者: Novi Quadrianto;V. Sharmanska;Oliver Thomas
DOI: 10.3389/frai.2021.612551
发表时间: 2021
期刊: Frontiers in artificial intelligence
影响因子: 4
作者: [Butcher B, Huang VS, Robinson C, Reffin J, Sgaier SK, Charles G, Quadrianto N]
通讯作者: Quadrianto N
Recycling Privileged Learning and Distribution Matching for Fairness
回收特权学习和分配匹配以实现公平
DOI: --
发表时间: 2017
期刊: Conference on Neural Information Processing Systems (NeurIPS -formerly NIPS)
影响因子: --
作者: [Quadrianto N]
通讯作者: Quadrianto N
DOI: 10.1609/aaai.v34i06.6572
发表时间: 2019-11
期刊:
影响因子: --
作者: [Artyom Gadetsky;Kirill Struminsky;Christopher Robinson;Novi Quadrianto;D. Vetrov]
通讯作者: Artyom Gadetsky;Kirill Struminsky;Christopher Robinson;Novi Quadrianto;D. Vetrov
海外基金