课题基金 / 基金详情

CAREER: New data integration approaches for efficient and robust meta-estimation, model fusion and transfer learning

CAREER: New data integration approaches for efficient and robust meta-estimation, model fusion and transfer learning
职业:新的数据集成方法,用于高效、稳健的元估计、模型融合和迁移学习
批准号:
2337943
负责人:
Emily Hector
金额:
$45.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-06-01 至 2029-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
统计科学的目的是通过从类似的实验观察中得出概括的结论来了解自然现象。随着最近的“大数据”和“开放科学”革命,科学家们已经将他们的重点从聚合单个观测数据转移到聚合大量公开可用的数据集。这一努力的前提是希望通过结合来自多个数据集的信息来提高研究结果的稳健性和普适性。例如,与仅基于一家医院的少数病例得出结论相比,将全美罕见疾病的结果数据结合起来可以描绘出更可靠的图景。同样,结合美国各地疾病风险因素的数据,可以区分地方和全国的健康趋势。到目前为止,对这些数据汇总目标的统计方法仅限于实际效用有限的简单环境。为了应对这一差距,该项目开发了从多个数据集聚合信息的新方法,涉及基于科学实践的三个不同的数据集成问题。所开发的方法直观、原则性强,对数据集之间的显著差异具有很强的稳健性,可广泛应用于医学、经济和社会科学等领域。在其他应用中,该项目将提供新的工具,从大型电子健康记录数据库中提取健康见解。该项目将支持本科生和研究生的培训、课程开发、招聘和对统计中任职人数不足的少数群体进行专业指导。此外,该项目将通过在服务不足的社区中的数据科学教师培训计划来影响STEM教育。该项目在三个基本的数据集成问题上开发了直观、原则性、健壮和高效的方法:元分析、模型融合和迁移学习。首先,该项目提供了一套荟萃分析方法,用于保护隐私-使用数据集相似性的新概念进行一次性估计和推理。该方法的主要创新之处在于对特定于数据集的参数和与经典元估计有一些相似之处的组合参数进行联合估计。其次,该项目建立了学习相似数据集的聚类的模型融合方法。这些方法的独特特征是模型融合,它沿着融合程度从高到低的频谱进行数据集成,从而不会强制集群数据集中的模型参数完全相等。第三,该项目开发了灵活和稳健的转移学习方法,利用历史信息提高感兴趣的目标数据集中的统计效率。这些方法的一个重要元素是适合源数据集的模型类型的灵活规范。这三套方法都重视推理输出的可解释性、统计效率和稳健性。该项目将三套拟议的方法统一在一个正式的数据集成框架下,该框架围绕数据集成的两个公理制定。数据整合理念渗透到收集数据的每一个科学研究领域,因此这项研究有助于医学、经济和社会科学等领域的科学努力。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Statistical science aims to learn about natural phenomena by drawing generalizable conclusions from an aggregate of similar experimental observations. With the recent “Big Data” and “Open Science” revolutions, scientists have shifted their focus from aggregating individual observations to aggregating massive publicly available datasets. This endeavor is premised on the hope of improving the robustness and generalizability of findings by combining information from multiple datasets. For example, combining data on rare disease outcomes across the United States can paint a more reliable picture than basing conclusions only on a small number of cases in one hospital. Similarly, combining data on disease risk factors across the United States can distinguish local from national health trends. To date, statistical approaches to these data aggregation objectives have been limited to simple settings with limited practical utility. In response to this gap, this project develops new methods for aggregating information from multiple datasets in three distinct data integration problems grounded in scientific practice. The developed approaches are intuitive, principled and robust to substantial differences between datasets, and are broadly applicable in medical, economic and social sciences, among others. Among other applications, the project will deliver new tools to extract health insights from large electronic health records databases. The project will support undergraduate and graduate student training, course development, and the recruitment and professional mentoring of under-represented minorities in statistics. Further, the project will impact STEM education through a data science teacher training program in underserved communities.This project develops intuitive, principled, robust and efficient methods in three essential data integration problems: meta-analysis, model fusion and transfer learning. First, the project delivers a set of meta-analysis methods for privacy-preserving one-shot estimation and inference using a new notion of dataset similarity. The primary novelty in the approach is the joint estimation of both dataset-specific parameters and a combined parameter that bears some similarity to the classic meta-estimator. Second, the project establishes model fusion methods that learn the clustering of similar datasets. The methods’ unique feature is a model fusion that dials data integration along a spectrum of more to less fusion and thereby does not force model parameters from clustered datasets to be exactly equal. Third, the project develops flexible and robust transfer learning approaches that leverage historical information for improved statistical efficiency in a target dataset of interest. An important element of these approaches is a flexible specification of the type of models fit to the source datasets. All three sets of methods place a premium on interpretability, statistical efficiency and robustness of the inferential output. The project unifies the three sets of proposed methods under a formal data integration framework formulated around two axioms of data integration. Data integration ideas pervade every field of scientific study in which data are collected, and so the research contributes to scientific endeavors in the medical, economic and social sciences, among others.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金