Computational Methods for Structured Data and Models
Computational Methods for Structured Data and Models
批准号:
2113079
负责人:
Maryclare Griffin
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-15 至 2024-07-31
中文摘要
许多领域的科学家,包括遗传学、神经科学、生态学和经济学,正在对更复杂的过程进行比以往任何时候都更丰富的测量。这提供了激动人心的机会来回答以前遥不可及的科学问题。例如,随着时间的推移,在许多地点收集的疾病流行率的测量提供了估计疾病传播的机会。潜变量模型将观测数据建模为未观测到的潜在随机变量的简单变换,是从复杂过程的测量中提取科学问题答案的流行方法。它们是灵活的,但它们也可能在计算上令人望而却步。因此,使用潜变量模型的科学家可能不得不满足于质量未知的近似,或者提供糟糕的估计或无法回答感兴趣的问题的临时简化。这个项目旨在开发新的方法来拟合潜变量模型,这些方法在计算上更高效、更可靠、更容易获得。参与该项目将培训各级统计人员,重点是统计研究中代表性不足人群的统计人员。具体地说,调查员将监督研究生参与研究,指导本科生参与外展材料的开发,参与当地高中的外展,并领导职业生涯早期教师的写作小组。当拟合潜变量模型时,出现了各种计算挑战。即使在表征潜变量模型的少量参数已知的情况下,也很难从给定观测数据的潜变量的条件分布中对其进行表征和模拟。此外,估计潜在变量模型的未知参数可能很困难,因为数据的可能性对应于高维积分,对于该高维积分,封闭形式的表达式可能不可用或难以评估。即使有可行的方法,如果没有开放源码软件和详细的教程,从业者也很难实施潜变量模型。因此,这个项目的目标是贡献(I)从高维潜在变量的条件分布模拟的新方法,(Ii)改进的潜在变量模型参数的最大似然估计方法,以及(Iii)允许从业者实施它们的通用统计软件。关于(I),PI计划开发新的路径方法,用于从高维潜在变量的条件分布模拟给定的数据,该数据利用目标条件分布与相关或近似分布的关系。关于(Ii),PI将开发用于潜在变量模型参数的最大似然估计的改进方法,该方法利用第一个目标中介绍的路径模拟方法。关于(Iii),PI将把新方法应用于疾病图谱和全基因组关联研究,并开发允许其他从业者实施这些方法的R包。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Scientists in many fields, including genetics, neuroscience, ecology, and economics, are obtaining richer measurements of more complex processes than ever before. This offers thrilling opportunities to answer scientific questions that were previously out of reach. For instance, measurements of disease prevalence collected at many locations over time provide the opportunity to estimate the spread of disease. Latent variable models, which model the observed data as a simple transformation of unobserved latent random variables, are a popular approach to extracting answers to scientific questions from measurements of complex processes. They are flexible, but they can also be computationally prohibitive. As a result, scientists using latent variable models may have to settle for approximations of unknown quality or ad-hoc simplifications that provide poor estimates or fail to answer questions of interest. This project aims to develop novel methods for fitting latent variable models that are more computationally efficient, reliable, and accessible. Involvement in the project will train statisticians at all levels, with a focus on statisticians from populations that are underrepresented in statistics research. Specifically, the investigator will supervise graduate student involvement in the research, guide the development of engaging outreach materials by undergraduate students, participate in outreach at local high schools, and lead writing groups for early-career faculty. A variety of computational challenges arise when fitting latent variable models. It can be difficult to characterize and simulate from the conditional distribution of the latent variables given observed data, even when the small number of parameters characterizing the latent variable model are known. Furthermore, it can be difficult to estimate the unknown parameters of a latent variable model because the likelihood of the data corresponds to a high dimensional integral for which a closed-form expression may be unavailable or expensive to evaluate. Even when feasible methods are available, it can be difficult for practitioners to implement latent variable models without access to open-source software and detailed tutorials. Accordingly, this project aims to contribute (i) novel methods for simulating from the conditional distributions of high dimensional latent variables, (ii) improved methods for maximum likelihood estimation of latent variable model parameters, and (iii) versatile statistical software that allows practitioners to implement them. Regarding (i), the PI plans to develop novel pathwise methods for simulating from the conditional distributions of high dimensional latent variables given data that leverage the relationship of the target conditional distribution to related or approximate distributions. Regarding (ii), the PI will develop improved methods for maximum likelihood estimation of latent variable model parameters that leverage the pathwise simulation methods introduced in the first aim. Regarding (iii), the PI will apply the new methods to disease mapping and genome-wide association studies and develop R packages that allow other practitioners to implement the methods.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhancing Underrepresented Participation in Mathematics & Statistics: Mentoring From Junior to Master’s
-
批准号:2130262
-
项目类别:Standard Grant
-
资助金额:$149.98万
-
财政年份:2022
-
负责人:Maryclare Griffin
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: