Empirical Process Theory for Complex Statistical Data Integration
Empirical Process Theory for Complex Statistical Data Integration
批准号:
2014971
负责人:
Takumi Saegusa
金额:
$20.44万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2024-06-30
中文摘要
如今,每个组织都从众多来源收集各种数据集。如果将这些数据集结合起来,提高推理质量将加速科学发现。然而,合并数据的统计分析具有挑战性,因为每个数据集通常只代表整个目标群体的一部分,并且因为合并的数据包含来自部分共享数据源的数据集的未识别的重复记录。本研究为解决由于合并数据的异质性和重复而导致的数据集成中不可避免的偏差问题提供了理论和方法基础。利用所提出的数据集成技术,将以前对较小人群的有限发现结合起来,推广到更广泛的人群。所提出的方法通过避免通过私人信息识别重复的记录链接,很好地保护了隐私。另一个好处是克服了单个数据源中相关信息的不足,而无需再次收集昂贵的(可能很小的)独立和相同分布的数据。该项目的预期成果将鼓励在现代数据分析中有效和社会地使用大量数据。研究生支持将用于跨学科活动和编写代码。该项目深入到经验过程理论,半和非参数推理,和抽样理论的交集。现有的理论和方法由于异质性和重复的存在,无法为研究具有偏倚和依赖性的复杂数据集成问题提供足够的工具。逆概率加权经验过程理论需要一种特殊的权重和变量独立结构。半参数和非参数推理通常依赖于独立和同分布样本的可用性。抽样理论处理特定设计中的依赖性,但侧重于参数模型,而不考虑有限总体框架中收集变量的随机性。为了解决概率工具和技术的缺乏,PI将开发一个统一的框架,与由多个框架调查驱动的加权经验过程相关联。这个加权的经验过程是可计算的,而不需要识别重复的选择。所提出的工具和技术将在研究一般样本选择和缺失数据机制方面发挥关键作用,例如方便样本,带有错误指定模型的半参数估计,以及重叠数据源中重复受试者的多次观测。正在研究的特殊问题包括(a)一般缺失机制下的一致极限定理,(b)数据集成模型错误规范下的稳健m估计,以及(c)集成对应于异构数据源的多个概率度量的一般理论。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Nowadays, every organization collects various data sets from numerous sources. If these data sets are combined, improved quality of inference will accelerate scientific discovery. Statistical analysis of merged data is, however, challenging because each data set often represents only a part of the entire target population and because combined data contain unidentified duplicated records from data sets which share data sources partially. This research provides theoretical and methodological foundations to address the issue of unavoidable bias in data integration arising from heterogeneity and duplication in merged data. With the proposed data integration technique, previously limited findings to smaller populations are combined to be generalized to a broader population. The proposed methodology serves well for privacy protection by avoiding record linkage that identifies duplication through private information. Another benefit is to overcome the shortage of relevant information in individual data sources without collecting costly(and possibly small) independent and identically distributed data all over again. Expected outcomes from this project will encourage the efficient and socially proper use of massive data in modern data analysis. The graduate student support will be used on interdisciplinary activities and writing codes. The project delves into the intersection of empirical process theory, semi- and non-parametric inference, and sampling theory. Existing theory and methods fail to provide sufficient tools to study complex data integration problems characterized by bias and dependence due to heterogeneity and duplication. Inverse probability-weighted empirical process theory requires a special independence structure on weights and variables. Semi- and non-parametric inference often relies on the availability of the independent and identically distributed sample. Sampling theory handles dependence in a specific design but focuses on a parametric model without accounting for randomness in collected variables in a finite population framework. To address the paucity of probabilistic tools and techniques, the PI will develop a unified framework in connection with a weighted empirical process motivated by multiple frame surveys. This weighted empirical process is computable without identifying duplicated selections. The proposed tools and techniques will play a critical role in studying a general sample selection and missing data mechanisms such as a convenience sample, semiparametric estimation with misspecified models, and multiple observations for duplicated subjects in overlapping data sources. The particular problems under investigation include (a) uniform limit theorems under general missingness mechanisms, (b) robust M-estimation under model misspecification for data integration, and (c) general theory to integrate multiple probability measures that correspond to heterogeneous data sources.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Semiparametric inference for merged data from multiple data sources
来自多个数据源的合并数据的半参数推理
DOI:
10.1016/j.jspi.2021.05.002
发表时间:
2022
期刊:
Journal of Statistical Planning and Inference
影响因子:
0.9
作者:
[Saegusa, Takumi]
通讯作者:
Saegusa, Takumi
Parametric Bootstrap Confidence Intervals for the Multivariate Fay–Herriot Model
多元 Fay–Herriot 模型的参数引导置信区间
DOI:
10.1093/jssam/smaa038
发表时间:
2022
期刊:
Journal of Survey Statistics and Methodology
影响因子:
2.1
作者:
[Saegusa, T.]
通讯作者:
Saegusa, T.
Mann–Whitney test for two‐phase stratified sampling
两相分层抽样的曼恩·惠特尼检验
DOI:
10.1002/sta4.321
发表时间:
2021
期刊:
Stat
影响因子:
1.7
作者:
[Saegusa, Takumi]
通讯作者:
Saegusa, Takumi
DOI:
10.1016/j.jspi.2021.05.001
发表时间:
2021
期刊:
Journal of Statistical Planning and Inference
影响因子:
0.9
作者:
[Saegusa, Takumi]
通讯作者:
Saegusa, Takumi
国内基金
海外基金
Neural Process模型的多样化高保真技术研究
-
批准号:62306326
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:王琦
-
依托单位:
磁转动超新星爆发中weak r-process的关键核反应
-
批准号:12375145
-
项目类别:面上项目
-
资助金额:52.00万元
-
批准年份:2023
-
负责人:金仕纶
-
依托单位:
多臂Bandit process中的Bayes非参数方法
-
批准号:71771089
-
项目类别:面上项目
-
资助金额:48.0万元
-
批准年份:2017
-
负责人:吴贤毅
-
依托单位: