CAREER: From Dirty Data to Fair Prediction: Data Preparation Framework for End-to-End Equitable Machine Learning
CAREER: From Dirty Data to Fair Prediction: Data Preparation Framework for End-to-End Equitable Machine Learning
批准号:
2341055
负责人:
Haewon Jeong
金额:
$55.82万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30
中文摘要
在一个人工智能正融入生活方方面面的时代,对尊重道德期望的人工智能系统的需求从未像现在这样重要。现代人工智能算法从例子中学习,创建更符合道德的系统应该从为算法提供更好的例子开始。虽然收集更多高质量的数据通常非常昂贵,但通过在数据准备阶段做出更好的选择来提高数据质量增加的额外成本最低。这项研究偏离了目前在培训阶段考虑道德目标的重点,培训阶段只是端到端数据科学生命周期的一小部分,而将数据准备渠道作为消除不必要的偏见和支持理想的道德目标的战略机会。该项目的教育范围包括提高人工智能学生和更广泛社区对伦理含义的理解,以及吸引和再培训人工智能领域的女性人才。该奖项围绕着一个关键问题:缺乏公平意识的数据准备产生的根本下游成本是什么,以及我们如何通过改进数据准备来实现端到端的公平?利用信息论的视角,PI将调查有偏见的信息如何通过数据准备管道从原始的肮脏数据流向干净的训练集,再流向经过训练的预测模型。具体地说,PI深入研究普遍存在的真实数据集问题,如缺失值、异构性和数据不平衡,以检查在处理这些问题时如何放大或减轻偏差。在这一分析的推动下,该项目将设计出公平意识的数据分配、数据编码和数据平衡技术,以更有效和高效地实现端到端的伦理目标。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In an era where AI is becoming integrated into every facet of life, the need for AI systems that respect ethical expectations has never been more crucial. Modern AI algorithms learn from examples, and the creation of more ethical systems should start by supplying better examples to the algorithm. While collecting more quality data is usually very expensive, improving the quality of data by making better choices in the data-preparation stage adds minimal extra cost. This research departs from the current focus of considering ethical goals in the training phase, which is merely a small part of the end-to-end data science lifecycle, and targets the data-preparation pipeline as a strategic opportunity for eliminating unwanted bias and bolstering desirable ethical objectives. The project’s education outreach includes enhancing the understanding of ethical implications among AI students and the wider community and attracting and retraining female talents in the AI field.This award centers around the critical question: What are the fundamental downstream costs arising from fairness-unaware data preparation and how can we move toward end-to-end fairness through improved data preparation? Employing an information-theoretic lens, the PI will investigate how biased information flows from the original, dirty data to the clean training set, to the trained prediction model through the data-preparation pipeline. Specifically, the PI delves into prevalent real-world dataset problems, such as missing values, heterogeneity, and data imbalance, to examine how bias can be either amplified or mitigated while handling these issues. Motivated by this analysis, the project will devise fairness-aware data imputation, data encoding, and data-balancing techniques that can attain end-to-end ethical goals more effectively and efficiently.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金