课题基金 / 基金详情

Latent Dependence and Identifiable, Graphical, Deep Modeling of Discrete Latent Variables

Latent Dependence and Identifiable, Graphical, Deep Modeling of Discrete Latent Variables
离散潜在变量的潜在依赖性和可识别、图形化、深度建模
批准号:
2210796
负责人:
Yuqi Gu
金额:
$15.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2025-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在数据科学时代,从教育到心理学再到医学,各个学科领域都出现了复杂的依赖和异构数据。潜在变量模型是处理此类复杂数据的强大统计方法。然而,现有的潜在变量分析的统计方法大多局限于相对简单的设置,不能满足现代高维应用的需要。例如,该项目的一个关键激励例子是个性化学习,教育工作者的目标是根据教育评估数据诊断个体学生在许多技能方面的潜在优势和劣势。在这种情况下,非常需要对学生的细粒度技能进行离散的统计诊断,了解各种潜在技能与潜在认知过程之间的关系,并制定有针对性的补救指导。为了实现这些目标,本项目旨在开发一套用于离散潜在变量建模的新统计工具。新的统计方法不仅适用于教育数据,也适用于心理学、医学、遗传学和健康科学的数据。这些工具将在公开可用的软件中实现。这些研究工具有望帮助从业者以统计原则的方式发现关于学生、患者和生物系统的隐藏信息。此外,本项目将为研究生和本科生提供多种培训机会,向他们介绍现代统计学中潜在变量模型的重要领域。本项目旨在推进离散潜变量建模的统计理论和方法,并提供适用于教育和其他应用的新颖统计算法。该项目有三个目标。首先是发展新的数学机制来研究具有潜在和图形成分的一般离散模型的可辨识性。这些技术将用于推导由教育科学驱动的模型的尖锐可识别条件。第二个目标是阐述两种新的具有离散潜在变量的生成模型:具有多层潜在结构的深度生成模型和编码硬层次潜在约束的概率图形模型。建立这些模型的可识别性,保证统计推断的有效性。由此产生的模型有望揭示几个应用中的潜在依赖关系,特别是与教育诊断和个性化学习相结合。第三个目标是开发新的可识别性假设检验、灵活的贝叶斯方法来同时推断潜在维数和其他参数,以及有效的结构学习程序来估计潜在的图形约束。该项目将在统计学、数据科学、心理学和教育科学的交叉领域为学员提供专业发展的机会。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the data science era, complex dependent and heterogeneous data emerge in various subject areas, from education to psychology to medicine. Latent variable models are powerful statistical approaches to tackle such complex data. However, existing statistical methods for analysis of latent variables are mostly limited to relatively simple settings and cannot meet the need for modern high dimensional applications. For example, one critical motivating example for this project is personalized learning, for which educators aim to diagnose individual students’ latent strengths and weaknesses across many skills based on educational assessment data. In this scenario, it is highly desirable to make discrete statistical diagnoses about student’s fine-grained skills, to understand the relationships between various latent skills and the underlying cognitive processes, and to develop targeted remedial instructions. To achieve these goals, this project aims to develop a suite of new statistical tools for discrete latent variable modeling. The new statistical methodology is intended to apply not only to educational data, but also to data from psychology, medicine, genetics, and health sciences. The tools will be implemented in publicly available software. These research tools are expected to help practitioners to uncover hidden information about students, patients, and biological systems in a statistically principled manner. In addition, this project will provide multiple training opportunities for graduate and undergraduate students, introducing them to the important area of latent variable models in modern statistics. This project aims to advance the statistical theory and methodology of discrete latent variable modeling and providing novel statistical algorithms applicable to education and other applications. The project has three objectives. The first is to develop new mathematical machinery to study identifiability in general discrete models with latent and graphical components. These techniques will be used to derive sharp identifiability conditions for models motivated by education sciences. The second objective is to elaborate two new families of generative models with discrete latent variables: deep generative models with multilayer latent structures, and probabilistic graphical models encoding hard hierarchical latent constraints. Identifiability of these models will be established, which would guarantee the validity of statistical inference. The resulting models are expected to shed light on latent dependencies in several applications, particularly, in conjunction with educational diagnoses and personalized learning. The third objective is to develop novel hypothesis testing of identifiability, flexible Bayesian methods to simultaneously infer latent dimensions and other parameters, and efficient structure learning procedures to estimate the latent graphical constraints. The project will offer opportunities for professional development of trainees at the interface of statistics, data science, psychology, and educational sciences.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021-09
期刊: J. Mach. Learn. Res.
影响因子: --
作者: [Yuqi Gu;E. Erosheva;Gongjun Xu;D. Dunson]
通讯作者: Yuqi Gu;E. Erosheva;Gongjun Xu;D. Dunson
Blessing of Dependence: Identifiability and Geometry of Discrete Models with Multiple Binary Latent Variables
依赖的祝福:具有多个二元潜变量的离散模型的可识别性和几何结构
DOI: --
发表时间: 2024
期刊: Bernoulli
影响因子: 1.5
作者: [Gu, Yuqi]
通讯作者: Gu, Yuqi
Bayesian Pyramids: identifiable multilayer discrete latent structure models for discrete data
贝叶斯金字塔:离散数据的可识别多层离散潜在结构模型
DOI: 10.1093/jrsssb/qkad010
发表时间: 2023
期刊: Journal of the Royal Statistical Society Series B: Statistical Methodology
影响因子: --
作者: [Gu, Yuqi, Dunson, David B]
通讯作者: Dunson, David B
Generic Identifiability of the DINA Model and Blessing of Latent Dependence
DINA 模型的通用可识别性和潜在依赖的祝福
DOI: 10.1007/s11336-022-09886-2
发表时间: 2022
期刊: Psychometrika
影响因子: 3
作者: [Gu, Yuqi]
通讯作者: Gu, Yuqi
共 6 条
    国内基金
    海外基金
    基于时间序列间分位相依性(quantile dependence)的风险值(Value-at-Risk)预测模型研究
    • 批准号:
      71903144
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      17.0万元
    • 批准年份:
      2019
    • 负责人:
      张申
    • 依托单位: