CAREER: Learning and Selecting Low-Dimensional Models from Incomplete Data
CAREER: Learning and Selecting Low-Dimensional Models from Incomplete Data
批准号:
2239479
负责人:
Daniel Pimentel-Alarcon
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-01 至 2028-01-31
中文摘要
大数据集通常有一个底层结构。确定这种结构可以根据一些变量预测感兴趣的结果,例如,根据药物的分子结构预测药物或疫苗的有效性。有各种各样的方法来学习数据集的底层结构并做出准确的预测。然而,当数据严重不完整时,就像许多现代数据集的情况一样,现有的方法始终无法识别数据的正确结构。更令人担忧的是,现有的方法无法验证所发现的结构是否正确。换句话说,只要数据不完整,任何现有方法学习到的结构都是不可信的,并且可能导致无法检测到的、任意错误的预测。该项目将(i)开发专门用于处理缺失数据的学习结构的方法,以及(ii)开发一种理论来验证通过任何方法(包括现有方法)学习的结构是否正确。反过来,这项研究将使科学家能够在药物发现、宏基因组学和机会筛选等大量应用中了解控制不完整数据集的结构,从而造福社会。此外,该项目将支持外联活动,通过实践活动、社交媒体活动、专题讨论会、课程和指导,让未被充分代表的少数民族参与本地和全国的机器学习。该项目的技术目标分为三个主要方面。第一个重点将研究一种将不完整数据映射到子空间的格拉斯曼流形的新方法,其中数据的底层结构可以通过求解由观测数据定义的舒伯特变量的约束优化来揭示。第二个重点将开发模型选择标准,以确定在候选结构集合中最适合不完整数据集的结构。这些标准将是赤池和贝叶斯信息标准和最小有效维的概括,适应于解释丢失的数据。这些标准将辅以拟合优度测试,以确定获胜的结构是否确实与数据很好地匹配。这些都是非常重要的任务,需要对丢失的数据进行特殊的考虑,丢失的数据可能会导致错误的结构适合任意大的数据集,其误差程度与正确的结构相同。最终,这一推动力的结果将决定来自特定结构的预测是否可信。第三个重点将在开源、易于使用的软件中实施我们的方法,以使更广泛的科学界受益,并在与我们正在进行的宏基因组学、单细胞测序、超声分类、细菌分类、药物发现和临床机会筛选等跨学科合作相关的数据集上进行测试。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Big datasets often have an underlying structure. Identifying such a structure allows predicting outcomes of interest based on a few variables, for example, predicting the effectiveness of a drug or vaccine based on the drug’s molecular structure. There exists a wide variety of methods to learn the underlying structure of a dataset and make accurate predictions. However, when data is severely incomplete, as is the case in many modern datasets, existing methods consistently fail to identify the correct structure of the data. More alarmingly, the existing methodology has no means to verify whether the structure found is correct or not. In other words, whenever data is incomplete, the structure learned by any existing method cannot be trusted and may result in undetectable, arbitrarily wrong predictions. This project will (i) develop methods to learn structures specifically tailored to handle missing data and (ii) develop a theory to verify whether the structure learned by any method (including existing ones) is correct or not. In turn, this research will enable scientists to learn the structures governing their incomplete datasets in a plethora of applications to the benefit of society, including drug discovery, metagenomics, and opportunistic screening. Furthermore, this project will support outreach activities to engage underrepresented minorities in machine learning, both locally and nationally, through hands-on activities, social media campaigns, symposia, courses, and mentoring.The technical aims of the project are divided into three main thrusts. The first thrust will investigate a new approach that maps incomplete data to the Grassmann manifold of subspaces, wherein the data’s underlying structure can be revealed by solving a constrained optimization over the Schubert varieties defined by the observed data. The second thrust will develop model-selection criteria to determine the structure that best fits an incomplete dataset, among a collection of candidate structures. These criteria will be generalizations of the Akaike and Bayes information criteria and the minimum effective dimension, adapted to account for missing data. These criteria will be complemented with a goodness-of-fit test to determine if the winning structure is, indeed, a good fit for the data. These are non-trivial tasks that require special considerations in light of missing data, which can consistently cause spurious structures fit arbitrarily large datasets with the same degree of error as the correct structures. Ultimately, the results from this thrust will allow determining whether the predictions stemming from a specific structure can be trusted or not. The third thrust will implement our methodology in open-source, easy-to-use software to benefit of the broader scientific community and test it on datasets related to our ongoing interdisciplinary collaborations in metagenomics, single-cell sequencing, sonotypes classification, bacteria classification, drug discovery, and clinical opportunistic screening.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: