CAREER: Mitigating the Lack of Labeled Training Data in Machine Learning Based on Multi-level Optimization
CAREER: Mitigating the Lack of Labeled Training Data in Machine Learning Based on Multi-level Optimization
批准号:
2339216
负责人:
Pengtao Xie
金额:
$50.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-09-01 至 2029-08-31
中文摘要
机器学习在自主驾驶、疾病早期检测、药物设计等众多应用中都取得了巨大的成功。机器学习模型的准确性在很大程度上依赖于大规模、人类标记的训练数据的可及性。然而,由于获取高级人类标签所涉及的高成本和数据隐私问题,在医疗保健、立法、环境科学等专业领域获取此类数据往往非常具有挑战性。该项目将通过提供算法、软件和系统来推动科学发展,这些算法、软件和系统可以自动生成高质量的标记数据,以缓解特定领域中缺乏标记训练数据的问题,并允许训练高度准确的机器学习模型。该项目将通过降低数据壁垒,显著扩大机器学习在各个应用领域的适用性,并将大幅降低人工数据标注的劳动力成本。例如,它将促进结构生物学和高能物理的科学发现,并简化无线通信的工程设计。它将有助于早期发现败血症、肺癌、帕金森氏病和睡眠呼吸暂停,改善患者的预后和生活质量。应用于化合物设计和水泥生产,所开发的技术有可能加快药物发现并降低能源消耗。为了实现创建高质量的标记训练数据的目标,该项目将开发基于多级优化和大型语言模型的三种新方法的互补范例,分别用于:1)端到端的标记数据生成;2)未标记数据的标注;以及3)特定于示例的改编/选择标记的源数据。首先,建议的数据生成方法将利用下游模型的最坏情况和特定类别的性能,为生成(具有复杂标签的)数据提供端到端和细粒度的指导,以提高下游模型的准确性和健壮性,并促进不同类别之间的平衡性能。其次,拟议的数据注释方法将利用端到端机制,该机制利用大型语言模型、一系列验证程序和可用的辅助信息,以最大限度地提高生成的标签的准确性。第三,所提出的适应/选择方法将区分目标域内部或外部的源实例,并随后端到端地确定特定于实例的适应/选择动作,以确保源数据的最佳使用。此外,提出的新的优化算法和分布式系统将有效地解决与多层次优化相关的新挑战,包括不可微性、与大型语言模型的优化器不兼容以及可伸缩性。该项目是第一个系统地利用多级优化来创建标签数据的项目,有效地解决了一个基本的知识鸿沟,即现有方法往往缺乏执行多个学习阶段的端到端执行的能力,因此在定制生成的数据以提高下游模型的性能方面存在不足。该项目的另一个重大创新是有效地利用大型语言模型进行数据标注,这将大大降低手动标注的成本。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning has demonstrated great success in numerous applications such as autonomous driving, early detection of diseases, drug design, etc. The accuracy of machine learning models highly depends on the accessibility of large-scale, human-labeled training data. However, such data is often very challenging to acquire in specialized domains such as healthcare, legislation, environmental sciences due to the high costs involved in obtaining high-grade human labels and data privacy concerns. This project will advance science by providing algorithms, software, and systems that can automatically generate high-quality labeled data to mitigate the lack of labeled training data in specific domains and and allow training of highly accurate machine learning models. The project will significantly broaden the applicability of machine learning across various application areas by lowering data barriers and will substantially reduce the labor costs of manual data annotation. For example, it will promote scientific discovery in structural biology and high-energy physics and streamline engineering design in wireless communication. It will facilitate the early detection of sepsis, lung cancer, Parkinson's disease, and sleep apnea, improving patient outcomes and quality of life. Applied to compound design and cement production, the developed technologies have the potential to expedite drug discovery and reduce energy consumption. To achieve the goal of creating high-quality labeled training data, this project will develop three complementary paradigms of novel approaches based on multi-level optimization and large language models, for: 1) end-to-end generation of labeled data; 2) annotation of unlabeled data; and, 3) example-specific adaptation/selection of labeled source data, respectively. First, the proposed data generation methods will leverage the worst-case and class-specific performance of downstream models to provide end-to-end and fine-grained guidance for generating data (with complex labels) that is tailored to improve the accuracy and robustness of downstream models, and to promote balanced performance across different classes. Second, the proposed data annotation methods will leverage an end-to-end mechanism that capitalizes on large language models, a sequence of verification procedures, and available side information to maximize the accuracy of generated labels. Third, the proposed adaptation/selection methods will distinguish between source examples that are inside or outside of a target domain and subsequently determine an example-specific adaptation/selection action end-to-end to ensure optimal use of source data. In addition, the proposed novel optimization algorithms and distributed systems will effectively tackle new challenges related to multi-level optimization, including non-differentiability, incompatibility with the optimizers of large language models, and scalability. This project represents the first one systematically leveraging multi-level optimization to create labeled data, effectively addressing a fundamental knowledge gap that existing methods often lack capabilities to perform end-to-end execution of multiple learning stages and therefore fall short in tailoring generated data to improve downstream models’ performance. Another significant innovation of this project is its effective harnessing of large language models for data annotation, which will substantially reduce the costs of manual labeling.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金