课题基金 / 基金详情

III: Small: Predictive Modeling from High-Dimensional, Sparsely and Irregularly Sampled, Longitudinal Data

III: Small: Predictive Modeling from High-Dimensional, Sparsely and Irregularly Sampled, Longitudinal Data
III:小:根据高维、稀疏和不规则采样的纵向数据进行预测建模
批准号:
2226025
负责人:
Vasant Honavar
金额:
$59.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-10-01 至 2025-09-30

项目摘要

项目成果

Vasant Honavar的其他基金

相似基金

相关文献

中文摘要
翻译
从一组个体随着时间的推移重复观察产生的纵向数据在许多应用中是常见的,包括健康科学、学习科学、社会科学、生命科学和经济学。这些数据提供了前所未有的机会来揭示某些测量变量(特征或协变量)随时间变化的模式与感兴趣的结果之间的关系,例如,经济崩溃、社会动荡、疾病发作、健康风险等。在现实世界中,变量的数量往往非常多;在任何给定的时间,往往只记录了变量的一小部分,导致数据稀少,遗漏观测的比例很高。此外,这些数据表现出复杂的相关性,如果没有得到适当的解释,可能会导致误导性的统计推断。更复杂的情况是,数据显示突然中断,这通常是由不可直接观察到的状态之间的转换(例如,从“健康”到“感染”)驱动的。大型数据集需要可伸缩的方法。在高风险的应用中,例如医疗保健,预测模型的人类可解释性是至关重要的。该项目将在可扩展机器学习方法方面取得实质性进展,用于从高维、不规则采样、稀疏、纵向健康数据的纵向结果预测建模。预测建模工具的开源实现将在许多领域找到应用,包括行为、社会、环境、经济、学习和健康科学。该项目将加强对数据科学和计算机科学(特别是人工智能)这两个具有重要国家意义的领域的不同研究生和本科生的研究性培训。与该项目相关的教育活动将有助于为不同的数据科学家、人工智能专家和健康科学、社会科学、学习科学及相关领域配备最先进的机器学习工具,用于从纵向数据进行预测建模。该项目将制作一个新的研究生课程和课程模块、样本项目等,涉及从纵向数据建立预测模型,并将其纳入数据科学课程。该项目将帮助来自不同背景的学生,包括女性和代表性不足的少数族裔,获得广泛的数据科学领域的教育、研究和职业机会。通过广泛传播所有研究成果(出版物、软件、数据集、课程材料),该项目的更广泛影响将进一步增强。该项目将开发一系列可扩展的深度核高斯过程回归算法,用于从具有复杂的先验未知关联结构的高维、稀疏和不规则的时间采样的纵向数据进行可解释的预测建模。由此产生的方法将能够发现未观察到的或隐藏的状态之间的转变模式,解释结果的突然中断。他们将能够通过学习数据所表现出的潜在的复杂的相关性结构来解释他们的预测,不仅通过识别驱动预测的变量,而且通过识别他们这样做的时间背景。该项目将使用模拟的纵向数据(具有不同的相关结构、不同的失配机制、不同的时间依赖变量重要性)、几个基准纵向数据集,以及最重要的是来自真实世界医疗保健应用程序的已确定的纵向电子健康记录数据和社会人口数据(与临床专家合作)对所产生的方法进行严格的经验性评估。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Longitudinal data resulting from repeated observations from a set of individuals over time are commonplace in many applications, including health sciences, learning sciences, social sciences, life sciences, and economics. Such data present unprecedented opportunities to uncover the relationship between the time- varying patterns of certain measured variables (features or covariates) and outcomes of interest e.g., economic meltdown societal unrest, disease onset, health risk, etc. In real-world settings, the number of variables is often very large; often only a small subset of variables is recorded at any given time, resulting in sparse data with a high proportion of missing observations. Furthermore, such data exhibit complex correlations which if not properly accounted for, can lead to misleading statistical inferences. Additional complications arise from the fact that the data exhibit abrupt discontinuities that are often driven by transitions between states that are not directly observable (e.g., from "healthy" to "infected"). Large size of data sets demand methods that are scalable. And in high stakes applications, e.g., healthcare, human interpretability of the predictive models is of paramount importance. The project will yield substantial advances over the current state-of-the-art in scalable machine learning methods for predictive modeling of longitudinal outcomes from high-dimensional, irregularly sampled, sparse, longitudinal health data. The open-source implementations of the predictive modeling tools will find applications in many domains including behavioral, social, environmental, economic, learning, and health sciences. The project will enhance the research-based training of a diverse graduate and undergraduate students in Data Sciences and Computer Science (especially Artificial Intelligence), areas of great national importance. The educational activities associated with the project will help equip a diverse cadre of Data Scientists, AI experts, and health sciences, social sciences, learning sciences, and related areas with state-of-the-art machine learning tools for predictive modeling from longitudinal data. The project will produce a new graduate course and course modules, sample projects, etc. on predictive modeling from longitudinal data to be integrated into Data Sciences curricula. The project will help introduce students from diverse backgrounds, including women and underrepresented minorities, to a broad range of educational, research, and career opportunities in Data Sciences. The broader impacts of the project will be further enhanced by broad dissemination of all research results (publications, software, data sets, course materials).The project will develop a family of scalable deep kernel gaussian process regression algorithms for interpretable predictive modeling from high dimensional, sparsely and irregularly time sampled, longitudinal data with complex, a priori unknown correlation structure. The resulting methods will be able to discover the patterns of transitions between unobserved or hidden states, account for abrupt discontinuities in outcomes. They will be able to explain their predictions by learning the underlying complex correlation structure exhibited by the data and by identifying not only the variables that drive the predictions, but also the temporal context in which they do so. The project will rigorously empirically evaluate the resulting methods with simulated longitudinal data (with different correlation structures, different missingness mechanisms, different time-dependent variable importance), several benchmark longitudinal data sets, and, most importantly, deidentified longitudinal electronic health records data and socio-demographic data from real-world healthcare applications (in collaboration with clinical experts).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
A Simple, Fast Algorithm for Continual Learning from High-Dimensional Data
一种简单、快速的高维数据持续学习算法
DOI: --
发表时间: 2023
期刊: Eleventh International Conference on Learning Representations
影响因子: --
作者: [Ashtekar, Neil, Honavar, Vasant G]
通讯作者: Honavar, Vasant G
Collaborative Research: RI: III: SHF: Small: Multi-Stakeholder Decision Making: Qualitative Preference Languages, Interactive Reasoning, and Explanation
AI Institute: Planning: Institute for AI-Enabled Materials Discovery, Design, and Synthesis
EAGER: Interpreting Black-Box Predictive Models Through Causal Attribution
BD Spokes: SPOKE: NORTHEAST: Collaborative Research: Integration of Environmental Factors and Causal Reasoning Approaches for Large-Scale Observational Health Research
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: