CAREER: From Dirty Data to Fair Prediction: Data Preparation Framework for End-to-End Equitable Machine Learning
CAREER: From Dirty Data to Fair Prediction: Data Preparation Framework for End-to-End Equitable Machine Learning
批准号:
2341055
负责人:
Haewon Jeong
金额:
$55.82万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30
中文摘要
在人工智能融入生活各个方面的时代,对尊重道德期望的人工智能系统的需求从未如此重要。现代人工智能算法从示例中学习,创建更多道德系统应该从为算法提供更好的示例开始。虽然收集更高质量的数据通常非常昂贵,但通过在数据准备阶段做出更好的选择来提高数据质量只会增加最小的额外成本。这项研究偏离了目前在培训阶段考虑道德目标的重点,这只是端到端数据科学生命周期的一小部分,并将数据准备管道作为消除不必要的偏见和支持理想道德目标的战略机会。该项目的教育推广包括提高AI学生和更广泛社区对伦理影响的理解,吸引和再培训AI领域的女性人才。该奖项围绕着一个关键问题:不公平的数据准备所产生的基本下游成本是什么,以及我们如何通过改进数据准备来实现端到端的公平?采用信息理论透镜,PI将调查有偏见的信息如何从原始的脏数据流向干净的训练集,通过数据准备管道流向训练的预测模型。具体来说,PI深入研究了普遍存在的现实数据集问题,如缺失值、异质性和数据不平衡,以研究在处理这些问题时如何放大或减轻偏倚。受此分析的启发,该项目将设计公平意识的数据估算,数据编码和数据平衡技术,可以更有效地实现端到端的道德目标。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估来支持。
英文摘要
In an era where AI is becoming integrated into every facet of life, the need for AI systems that respect ethical expectations has never been more crucial. Modern AI algorithms learn from examples, and the creation of more ethical systems should start by supplying better examples to the algorithm. While collecting more quality data is usually very expensive, improving the quality of data by making better choices in the data-preparation stage adds minimal extra cost. This research departs from the current focus of considering ethical goals in the training phase, which is merely a small part of the end-to-end data science lifecycle, and targets the data-preparation pipeline as a strategic opportunity for eliminating unwanted bias and bolstering desirable ethical objectives. The project’s education outreach includes enhancing the understanding of ethical implications among AI students and the wider community and attracting and retraining female talents in the AI field.This award centers around the critical question: What are the fundamental downstream costs arising from fairness-unaware data preparation and how can we move toward end-to-end fairness through improved data preparation? Employing an information-theoretic lens, the PI will investigate how biased information flows from the original, dirty data to the clean training set, to the trained prediction model through the data-preparation pipeline. Specifically, the PI delves into prevalent real-world dataset problems, such as missing values, heterogeneity, and data imbalance, to examine how bias can be either amplified or mitigated while handling these issues. Motivated by this analysis, the project will devise fairness-aware data imputation, data encoding, and data-balancing techniques that can attain end-to-end ethical goals more effectively and efficiently.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金