课题基金 / 基金详情

Improving Causal Inference Methods in Statistics for Analyzing Big Data

Improving Causal Inference Methods in Statistics for Analyzing Big Data
改进统计学中用于分析大数据的因果推理方法
批准号:
RGPIN-2018-05044
负责人:
Karim, Mohammad
金额:
$1.53万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Karim, Mohammad的其他基金

相似基金

相关文献

中文摘要
翻译
计算机器的可用性不断提高,成本不断下降,智能和基于云的技术的广泛应用,导致了为商业、实用和科学目的收集大规模信息的趋势日益增长。这些数据库通常包含相当多的变量,覆盖了大量的人群,随访时间长,与那些来自精心控制的随机实验相比,更好地反映了“现实世界”的日常实践。然而,这些数据集的收集主要不是为了研究目的,并且在缺乏随机化的情况下,混淆对探索结果与干预之间的因果关系构成了重大挑战。在统计和因果推理文献中有大量关于混杂调整的文献,这些文献指导我们选择合适的变量进行调整和控制,例如控制混杂因素和风险因素,但不调整仪器和噪声变量。由于这些具有数千个变量的数据库的复杂性和庞大规模,领域专家无法(i)手工挑选重要的混杂因素或确定哪些变量是工具,(ii)合理正确地猜测干预模型(在倾向得分上下文中)或结果模型中协变量的功能形式,(iii)充分评估这么多变量的协变量平衡。******为了应对这些挑战,本提案中有四个具体的研究目标。1. 结合因果推理文献中建立的原则,在高维环境中发展混杂选择方法。2. 研究各种数据自适应方法在高维模型错配情况下的鲁棒性。3. 在从高维协变量估计的倾向得分的背景下,提出适当的指标来评估“协变量平衡”。4. 在有纵向数据的情况下调查上述问题。这些方法将通过理论发展、现实应用和现实模拟进行评估。******我处于一个独特的跨学科研究环境中,作为UBC医学院的助理教授,圣保罗医院的生物统计学家,UBC统计系的校友,与UBC和麦吉尔有着密切的研究联系。在这个大数据时代,对受过统计建模训练的学生有着巨大的需求,他们可以在分析大数据集时考虑因果结构。在跨学科的环境中培养高素质的人才是这项研究的重要组成部分。受训者将接受培训并获得高质量的研究数据集以及将对现实生活产生影响的方法和应用研究问题。
英文摘要
The increasing availability, declining cost of computational machineries and wider application of smart and cloud-based technologies have led to a growing trend of collecting large-scale information for business, utilitarian and scientific purposes. These databases generally contain a considerable number of variables, cover substantially large populations with long follow-up, and better reflect ‘real-world' daily practices compared to those derived from carefully controlled randomized experiments. However, these datasets are not primarily collected for research purposes, and in the absence of randomization, confounding poses a critical challenge in exploring the cause-and-effect relationship between the outcome and the intervention. There is a vast literature on confounding adjustment in the statistical and causal inference literature that guides us to select appropriate variables to adjust and control, e.g., controlling for confounders and risk factors, but not adjusting for instruments and noise variables. Due to the complexity and large size of these databases with thousands of variables, it is not tenable for a domain expert to (i) hand-pick the important confounders or identify which variables are instruments, (ii) reasonably correctly guess the functional form of the covariates in the intervention model (in the propensity score context) or the outcome model, (iii) adequately assess the covariate balance for so many variables. ******To address these challenges, there are four specific research objectives in this proposal. 1. To develop confounder selection approaches in a high dimensional setting incorporating the principles established in the causal inference literature. 2. To study the robustness of various data-adaptive methods in the context of model misspecification in a high dimensional setting. 3. To propose appropriate metrics for assessing the ‘covariate balance' in the context of propensity scores estimated from high-dimensional covariates. 4. To investigate the above issues when longitudinal data are available. These methods will be evaluated through theoretical developments, real-life applications, and via realistic simulations. ******I am positioned in a unique interdisciplinary research environment, as an Assistant Professor in the UBC Faculty of Medicine, a biostatistician at St. Paul's hospital, and an alumnus from the Statistics department, UBC, with close research ties with UBC and McGill. In this big-data era, there are huge demands for students with training in statistical modeling who can take causal structures into consideration while analyzing a large data set. Training of highly qualified personnel within an interdisciplinary environment is an essential component of this research. Trainees will receive training and access to high-quality research datasets and methodological and applied research questions that will have a real-life impact.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2022
  • 负责人:
    Karim, Mohammad
  • 依托单位:
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2021
  • 负责人:
    Karim, Mohammad
  • 依托单位:
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2020
  • 负责人:
    Karim, Mohammad
  • 依托单位:
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2019
  • 负责人:
    Karim, Mohammad
  • 依托单位:
海外基金