课题基金 / 基金详情

CRII: CIF: Information Theoretic Measures for Fairness-aware Supervised Learning

CRII: CIF: Information Theoretic Measures for Fairness-aware Supervised Learning
CRII:CIF:公平意识监督学习的信息论措施
批准号:
2246058
负责人:
Mohamed Nafea
金额:
$17.49万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2025-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
尽管机器学习(ML)系统在完成复杂任务方面取得了越来越大的成功,但它们在制定或辅助影响人们生活的重大决策(例如,大学录取、医疗保健、预测性警务)方面的应用越来越多,引发了人们对潜在歧视行为的担忧。机器学习系统中的不公平结果是由用于训练它们的数据中的历史偏差造成的。仅仅为了最小化预测误差而设计的学习算法可能会继承甚至加剧这种偏差;特别是当观察到的对产生准确决策至关重要的个人属性由于现有的社会和文化不平等而受到其群体身份(例如种族或性别)的偏见时。在数据层面理解和衡量这些偏见是一个具有挑战性但又至关重要的问题,它可以为消除数据偏见和调整学习系统以最大限度地减少歧视提供建设性的见解和方法,并提出对政策变革和基础设施发展的需求。该项目旨在建立一个全面的框架,以精确量化个人属性对决策准确性和不公平性的边际影响,使用信息和博弈论以及因果推理的工具,以及法律和社会科学对公平的定义。这项多学科的努力将为公平数据驱动自动化系统领域的从业者提供指导和设计见解,并为有关人工智能社会后果的公开辩论提供信息。以前的大部分工作是从学习算法的角度出发,通过对学习者的结果施加统计或反事实公平性约束,并设计满足该约束的学习者来阐述算法公平性问题。由于所考虑的公平性问题源于有偏见的数据,仅向预测任务添加约束可能无法提供其基本局限性的整体视图。这个项目从不同的角度看待公平问题,而不是问“对于给定的学习者,我们如何实现公平”?,它要求“对于给定的数据集,数据中固有的权衡是什么,并基于这些,我们可以设计的最佳学习器是什么”?在监督学习模型中,所提出问题的挑战在于个体属性(协变量)、群体身份(受保护特征)、目标变量(标签)和预测结果(决策)之间的关联/因果关系结构复杂。在分析数据集时,通过精心设计的衡量变量之间复杂的相关性/因果关系结构以及准确性和公平性目标之间的内在紧张关系,从数据中量化协变量对决策准确性和歧视的边际影响。随后,将研究如何利用量化影响来指导下游机器学习系统,以提高其可实现的准确性-公平性权衡。重要的是,提出的框架提供了可解释的解决方案,其中在学习系统中包含某些属性是通过它们对准确和公平决策的重要性来解释的。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Despite the growing success of Machine Learning (ML) systems in accomplishing complex tasks, their increasing use in making or aiding consequential decisions that affect people’s lives (e.g., university admission, healthcare, predictive policing) raises concerns about potential discriminatory practices. Unfair outcomes in ML systems result from historical biases in the data used to train them. A learning algorithm designed merely to minimize prediction error may inherit or even exacerbate such biases; particularly when observed attributes of individuals, critical for generating accurate decisions, are biased by their group identities (e.g., race or gender) due to existing social and cultural inequalities. Understanding and measuring these biases-- at the data level-- is a challenging yet crucial problem, leading to constructive insights and methodologies for debiasing the data and adapting the learning system to minimize discrimination, as well as raising the need for policy changes and infrastructural development. This project aims to establish a comprehensive framework for precisely quantifying the marginal impact of individuals’ attributes on accuracy and unfairness of decisions, using tools from information and game theories and causal inference, along with legal and social science definitions of fairness. This multi-disciplinary effort will provide guidelines and design insights for practitioners in the field of fair data-driven automated systems and inform the public debate on social consequences of artificial intelligence.The majority of previous work formulates the algorithmic fairness problem from the viewpoint of the learning algorithm by enforcing a statistical or counterfactual fairness constraint on the learner’s outcome and designing a learner that meets it. As the considered fairness problem originates from biased data, merely adding constraints to the prediction task might not provide a holistic view of its fundamental limitations. This project looks at the fairness problem through different lens, where instead of asking “for a given learner, how can we achieve fairness”?, it asks “for a given dataset, what are the inherent tradeoffs in the data, and based on these, what is the best learner we can design”?. In supervised learning models, the challenge in the proposed problem lies in the complex structures of correlation/causation among individuals’ attributes (covariates), their group identities (protected features), the target variable (label), and the prediction outcome (decision). In analyzing the dataset, the marginal impacts of covariates on the accuracy and discrimination of decisions are quantified from the data, via carefully designed measures accounting for the complex correlation/causation structures among variables and the inherent tension between accuracy and fairness objectives. Subsequently, methods to exploit the quantified impacts in guiding downstream ML systems to improve their achievable accuracy-fairness tradeoff will be investigated. Importantly, the proposed framework provides explainable solutions, where the inclusion of certain attributes in the learning system is explained by their importance for accurate as well as fair decisions.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Wolbachia的cif因子与天麻蚜蝇dsx基因协同调控生殖不育的机制研究
  • 批准号:
    JCZRQN202501187
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
  • 依托单位:
SHR和CIF协同调控植物根系凯氏带形成的机制
  • 批准号:
    31900169
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2019
  • 负责人:
    李朋雪
  • 依托单位: