课题基金 / 基金详情

CAREER: Causal Modeling for Data Quality and Bias Mitigation

CAREER: Causal Modeling for Data Quality and Bias Mitigation
职业:数据质量和偏差缓解的因果建模
批准号:
2340124
负责人:
Babak Salimi
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目提出了一种受数据库方法学启发的新方法,以解决算法系统中的偏见的重大挑战,特别是在信用评分,医疗诊断,预测性警务和刑事司法系统等敏感领域。 通过认识到这种偏见往往源于基础数据,该倡议将算法偏见重新定义为数据质量管理问题。该项目强调数据质量管理的关键方面,如准确性、完整性和一致性,旨在开发可显著增强这些系统的可信度和社会影响的方法。通过使用这些基本的数据质量原则进行因果建模,它采取了一种战略方法来识别和解决算法偏差的根本原因。这一努力不仅标志着数据科学领域的重大进步,而且通过倡导公平,准确和可靠的决策过程,为国家和公共福利做出了重大贡献,从而全面促进国家健康,繁荣和福祉。 该计划设想通过各种跨学科座谈会,研讨会和课外学习机会广泛传播其动机,方法和文物。该项目通过四种方法解决算法偏差:1)开发新的可扩展的数据修复算法,旨在修复有关特殊类别完整性约束的数据,这些约束可以捕获用于训练机器学习(ML)模型的数据的统计细微差别。2)建立一个全面的数据去偏见框架,能够解决各种数据偏见和质量问题。3)实施方法来量化算法决策中的不确定性,特别是基于ML模型,其中不确定性源于偏差和数据质量问题,由于信息不完整而无法完全恢复和删除。4)最后,该项目侧重于开发根本原因分析方法,以确定动态数据环境中的潜在问题和自适应去偏置,并在数据处理管道中纳入主动干预措施,以持续缓解偏置。这一多方面的战略旨在推进数据质量管理、ML数据清洗和负责任的数据科学领域,显著提高数据驱动决策系统的可靠性、公平性和准确性。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project presents a novel approach, inspired by database methodologies, to address the significant challenge of bias in algorithmic systems, particularly in sensitive domains such as credit scoring, medical diagnostics, predictive policing, and the criminal justice system. By recognizing that such biases often stem from the underlying data, the initiative redefines algorithmic bias as a data quality management issue. Emphasizing critical aspects of data quality management such as accuracy, completeness, and consistency, the project aims to develop methods that significantly enhance the trustworthiness and societal impact of these systems. Incorporating causal modeling with these essential data quality principles, it takes a strategic approach to identifying and addressing the root causes of algorithmic bias. This effort not only marks a significant advancement in the field of data science but also contributes substantially to national and public welfare by advocating for decision-making processes that are fair, accurate, and reliable, thereby promoting national health, prosperity, and well-being in a comprehensive manner. This plan envisions a wide-ranging dissemination of its motivation, approach, and artifacts through a diverse array of interdisciplinary colloquia, seminars, and co-curricular learning opportunities. This project addresses algorithmic bias through a fourfold approach: 1) Developing new, scalable algorithms for data repair, designed for repairing data concerning a special class of integrity constraints that can capture the statistical nuances of data used for training machine learning (ML) models. 2) Establishing a holistic data debiasing framework capable of addressing various data biases and quality issues. 3) Implementing methods to quantify uncertainty in algorithmic decision-making, particularly based on ML models, where the uncertainty stems from bias and data quality issues that cannot be fully recovered and removed due to incomplete information. 4) Lastly, the project focuses on developing methods for root-cause analysis to identify underlying issues and adaptive debiasing in dynamic data environments, incorporating proactive interventions in data processing pipelines for ongoing bias mitigation. This multifaceted strategy aims to advance the fields of data quality management, data cleaning for ML, and responsible data science, significantly enhancing the reliability, fairness, and accuracy of data-driven decision-making systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金