SLES: Foundations of Safety-Aware Learning in the Wild
SLES: Foundations of Safety-Aware Learning in the Wild
批准号:
2331669
负责人:
Sharon Li
金额:
$79.31万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-01-01 至 2026-12-31
中文摘要
今天的机器学习(ML)模型必须在日益动态和不可预测的环境中运行。部署在野外的模型面临的一个关键挑战是,除了已知的分布内(ID)数据外,它们还将遇到未知的分布外(OOD)数据。目前用于训练ML模型的方法,特别是在有监督的设置中,众所周知是脆弱的,并且缺乏必要的安全意识,例如,OOD数据可能被盲目地分类为高置信度的已知类别。该项目的创新之处在于开发了具有安全意识的学习方法和理论保证,当模型部署在野外时,这些方法和理论保证可以被证实地检测OOD数据。该项目的影响是提高依赖于人工智能(AI)分类的广泛下游应用的安全性,包括交通、医疗、商业和科学发现,以便它们能够正确处理意外输入。该项目将使AI更好地理解它知道和不知道的东西,从而避免意外输入,而不是以最高的信心错误分类。从技术上讲,研究团队将设计新的算法,以利用在模型的部署环境中随处可见的大量真实世界的未标记数据。*从这类数据中学习可能具有挑战性,因为它具有异构性(与ID和OOD数据混合)和非平稳性(随时间变化)。为了应对这些挑战,该项目设计了新的机器学习算法,可证明使用未标记的野生数据进行OOD安全感知学习、在线OOD检测以适应不断变化的环境,以及使用基础模型进行OOD检测。学习框架将通过经验实验和理论理解的平衡进行评估。这项研究得到了国家科学基金会和开放慈善机构之间的合作伙伴关系的支持。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning (ML) models today must operate amid increasingly dynamic and unpredictable environments. A crucial challenge for models deployed in the wild is that they will encounter unknown out-of-distribution (OOD) data, in addition to the known in-distribution (ID) data. Current approaches for training ML models, particularly in a supervised setting, are known to be brittle and lack necessary safety awareness, e.g., OOD data may be blindly classified as a known class with high confidence. The project's novelties are developing safety-aware learning methodologies and theoretical guarantees that can provably detect OOD data as models are deployed in the wild. The project's impacts are to enhance safety for a broad range of downstream applications that depend on artificial intelligence (AI) classification, including transportation, healthcare, commerce, and scientific discovery, so that they can properly handle unexpected input.This project will make AI understand better what it knows and doesn't know, so that it abstains from unexpected input instead of wrongly classifying with supreme confidence. Technically, the team of researchers will design new algorithms that can leverage a large amount of real-world unlabeled data that arises ubiquitously in the model's deployment environment. Learning from such data can be challenging due to its heterogeneity (mixed with ID and OOD data) and non-stationarity (changes over time). To address the challenges, the project designs new machine learning algorithms that provably use unlabeled wild data for OOD safety-aware learning, online OOD detection to adapt to changing environments, and OOD detection with foundation models. The learning framework will be evaluated by a balance of empirical experimentation and theoretical understanding.This research is supported by a partnership between the National Science Foundation and Open Philanthropy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Foundations of Human-Centered Machine Learning in the Wild
-
批准号:2237037
-
项目类别:Continuing Grant
-
资助金额:$59.93万
-
财政年份:2023
-
负责人:Sharon Li
-
依托单位:
海外基金