CAREER: From Dirty Data to Fair Prediction: Data Preparation Framework for End-to-End Equitable Machine Learning
CAREER: From Dirty Data to Fair Prediction: Data Preparation Framework for End-to-End Equitable Machine Learning
批准号:
2341055
负责人:
Haewon Jeong
金额:
$55.82万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30
中文摘要
在一个人工智能正融入生活方方面面的时代,对尊重道德期望的人工智能系统的需求从未像现在这样重要。现代人工智能算法从例子中学习,而创建更道德的系统应该从为算法提供更好的例子开始。虽然收集更多高质量的数据通常非常昂贵,但通过在数据准备阶段做出更好的选择来提高数据质量只会增加最小的额外成本。本研究偏离了目前在培训阶段考虑道德目标的重点,这只是端到端数据科学生命周期的一小部分,并将数据准备管道作为消除不必要偏见和支持理想道德目标的战略机会。该项目的教育外展包括提高人工智能学生和更广泛社区对伦理影响的理解,以及吸引和再培训人工智能领域的女性人才。该奖项围绕着一个关键问题:不考虑公平性的数据准备所产生的基本下游成本是什么?我们如何通过改进数据准备来实现端到端的公平性?PI将采用信息论的视角,研究有偏差的信息如何从原始的脏数据流向干净的训练集,再通过数据准备管道流向训练好的预测模型。具体来说,PI深入研究了现实世界中普遍存在的数据集问题,如缺失值、异质性和数据不平衡,以检查在处理这些问题时如何放大或减轻偏见。在此分析的推动下,该项目将设计出具有公平性意识的数据输入、数据编码和数据平衡技术,这些技术可以更有效地实现端到端的道德目标。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In an era where AI is becoming integrated into every facet of life, the need for AI systems that respect ethical expectations has never been more crucial. Modern AI algorithms learn from examples, and the creation of more ethical systems should start by supplying better examples to the algorithm. While collecting more quality data is usually very expensive, improving the quality of data by making better choices in the data-preparation stage adds minimal extra cost. This research departs from the current focus of considering ethical goals in the training phase, which is merely a small part of the end-to-end data science lifecycle, and targets the data-preparation pipeline as a strategic opportunity for eliminating unwanted bias and bolstering desirable ethical objectives. The project’s education outreach includes enhancing the understanding of ethical implications among AI students and the wider community and attracting and retraining female talents in the AI field.This award centers around the critical question: What are the fundamental downstream costs arising from fairness-unaware data preparation and how can we move toward end-to-end fairness through improved data preparation? Employing an information-theoretic lens, the PI will investigate how biased information flows from the original, dirty data to the clean training set, to the trained prediction model through the data-preparation pipeline. Specifically, the PI delves into prevalent real-world dataset problems, such as missing values, heterogeneity, and data imbalance, to examine how bias can be either amplified or mitigated while handling these issues. Motivated by this analysis, the project will devise fairness-aware data imputation, data encoding, and data-balancing techniques that can attain end-to-end ethical goals more effectively and efficiently.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金