CAREER: New data integration approaches for efficient and robust meta-estimation, model fusion and transfer learning
CAREER: New data integration approaches for efficient and robust meta-estimation, model fusion and transfer learning
批准号:
2337943
负责人:
Emily Hector
金额:
$45.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-06-01 至 2029-05-31
中文摘要
统计科学旨在通过从大量相似的实验观察中得出可概括的结论来了解自然现象。随着最近的“大数据”和“开放科学”革命,科学家们已经将他们的重点从汇总个人观察转移到汇总大量公开可用的数据集。这一努力的前提是希望通过结合来自多个数据集的信息来提高研究结果的稳健性和泛化性。例如,与仅根据一家医院的少数病例得出结论相比,综合美国各地罕见疾病结果的数据可以描绘出更可靠的图景。同样,结合美国各地疾病风险因素的数据可以区分地方和国家的健康趋势。迄今为止,实现这些数据汇总目标的统计方法仅限于简单的设置,实际效用有限。为了弥补这一差距,本项目开发了新的方法,用于在基于科学实践的三个不同的数据集成问题中从多个数据集聚合信息。所开发的方法直观、有原则,对数据集之间的重大差异具有稳健性,广泛适用于医学、经济和社会科学等领域。在其他应用程序中,该项目将提供从大型电子健康记录数据库中提取健康信息的新工具。该项目将支持本科生和研究生培训、课程开发以及统计领域代表性不足的少数民族的招聘和专业指导。此外,该项目将通过数据科学教师培训计划在服务不足的社区影响STEM教育。本项目在元分析、模型融合和迁移学习这三个基本数据集成问题上开发了直观、有原则、稳健和高效的方法。首先,该项目提供了一套元分析方法,用于使用数据集相似度的新概念进行保护隐私的一次性估计和推断。该方法的主要新颖之处在于对数据集特定参数和与经典元估计器有一些相似之处的组合参数进行联合估计。其次,建立了学习相似数据集聚类的模型融合方法。该方法的独特之处在于模型融合,它将数据沿多融合到少融合的频谱进行整合,因此不会强制来自聚类数据集的模型参数完全相等。第三,该项目开发了灵活而强大的迁移学习方法,利用历史信息来提高目标数据集的统计效率。这些方法的一个重要元素是对适合源数据集的模型类型的灵活规范。这三种方法都注重可解释性、统计效率和推断输出的稳健性。该项目将提出的三组方法统一在一个正式的数据集成框架下,该框架围绕两个数据集成公理制定。数据集成思想渗透到收集数据的每一个科学研究领域,因此数据集成研究有助于医学、经济和社会科学等领域的科学努力。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Statistical science aims to learn about natural phenomena by drawing generalizable conclusions from an aggregate of similar experimental observations. With the recent “Big Data” and “Open Science” revolutions, scientists have shifted their focus from aggregating individual observations to aggregating massive publicly available datasets. This endeavor is premised on the hope of improving the robustness and generalizability of findings by combining information from multiple datasets. For example, combining data on rare disease outcomes across the United States can paint a more reliable picture than basing conclusions only on a small number of cases in one hospital. Similarly, combining data on disease risk factors across the United States can distinguish local from national health trends. To date, statistical approaches to these data aggregation objectives have been limited to simple settings with limited practical utility. In response to this gap, this project develops new methods for aggregating information from multiple datasets in three distinct data integration problems grounded in scientific practice. The developed approaches are intuitive, principled and robust to substantial differences between datasets, and are broadly applicable in medical, economic and social sciences, among others. Among other applications, the project will deliver new tools to extract health insights from large electronic health records databases. The project will support undergraduate and graduate student training, course development, and the recruitment and professional mentoring of under-represented minorities in statistics. Further, the project will impact STEM education through a data science teacher training program in underserved communities.This project develops intuitive, principled, robust and efficient methods in three essential data integration problems: meta-analysis, model fusion and transfer learning. First, the project delivers a set of meta-analysis methods for privacy-preserving one-shot estimation and inference using a new notion of dataset similarity. The primary novelty in the approach is the joint estimation of both dataset-specific parameters and a combined parameter that bears some similarity to the classic meta-estimator. Second, the project establishes model fusion methods that learn the clustering of similar datasets. The methods’ unique feature is a model fusion that dials data integration along a spectrum of more to less fusion and thereby does not force model parameters from clustered datasets to be exactly equal. Third, the project develops flexible and robust transfer learning approaches that leverage historical information for improved statistical efficiency in a target dataset of interest. An important element of these approaches is a flexible specification of the type of models fit to the source datasets. All three sets of methods place a premium on interpretability, statistical efficiency and robustness of the inferential output. The project unifies the three sets of proposed methods under a formal data integration framework formulated around two axioms of data integration. Data integration ideas pervade every field of scientific study in which data are collected, and so the research contributes to scientific endeavors in the medical, economic and social sciences, among others.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金