课题基金 / 基金详情

Improving Causal Inference Methods in Statistics for Analyzing Big Data

Improving Causal Inference Methods in Statistics for Analyzing Big Data
改进统计学中用于分析大数据的因果推理方法
批准号:
RGPIN-2018-05044
负责人:
Karim, Mohammad
金额:
$1.53万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Karim, Mohammad的其他基金

相似基金

相关文献

中文摘要
翻译
计算机器的可用性不断增加,成本不断下降,智能和云技术的应用越来越广泛,导致为商业、实用和科学目的收集大规模信息的趋势日益增长。这些数据库通常包含相当数量的变量,覆盖了大量的长期随访人群,与精心控制的随机实验相比,更好地反映了“真实世界”的日常实践。然而,这些数据集主要不是为了研究目的而收集的,在没有随机化的情况下,混淆对探索结果和干预之间的因果关系构成了关键挑战。在统计和因果推理文献中有大量关于混杂调整的文献,指导我们选择适当的变量进行调整和控制,例如,控制混杂因素和风险因素,但不调整仪器和噪声变量。由于这些数据库的复杂性和庞大的规模,成千上万的变量,这是站不住脚的领域专家(i)手工挑选重要的混杂因素或确定哪些变量是工具,(ii)合理正确地猜测干预模型(在倾向评分的背景下)或结果模型中的协变量的函数形式,(iii)充分评估协变量平衡这么多的变量。* 为了应对这些挑战,本提案中有四个具体的研究目标。1.结合因果推理文献中建立的原则,在高维环境中开发混杂因素选择方法。2.研究在高维环境中模型误设定的背景下各种数据自适应方法的鲁棒性。3.提出适当的度量标准,用于评估从高维协变量估计的倾向评分背景下的“协变量平衡”。4.在纵向数据可用的情况下,调查上述问题。这些方法将通过理论发展,实际应用和现实模拟进行评估。** 我被定位在一个独特的跨学科研究环境中,作为UBC医学院的助理教授,圣保罗医院的生物统计学家,以及UBC统计系的校友,与UBC和麦吉尔大学有着密切的研究联系。在这个大数据时代,对受过统计建模培训的学生有巨大的需求,他们可以在分析大型数据集时考虑因果结构。在跨学科环境中培养高素质人才是这项研究的重要组成部分。学员将接受培训,并获得高质量的研究数据集和方法和应用研究问题,将有现实生活中的影响。
英文摘要
The increasing availability, declining cost of computational machineries and wider application of smart and cloud-based technologies have led to a growing trend of collecting large-scale information for business, utilitarian and scientific purposes. These databases generally contain a considerable number of variables, cover substantially large populations with long follow-up, and better reflect ‘real-world' daily practices compared to those derived from carefully controlled randomized experiments. However, these datasets are not primarily collected for research purposes, and in the absence of randomization, confounding poses a critical challenge in exploring the cause-and-effect relationship between the outcome and the intervention. There is a vast literature on confounding adjustment in the statistical and causal inference literature that guides us to select appropriate variables to adjust and control, e.g., controlling for confounders and risk factors, but not adjusting for instruments and noise variables. Due to the complexity and large size of these databases with thousands of variables, it is not tenable for a domain expert to (i) hand-pick the important confounders or identify which variables are instruments, (ii) reasonably correctly guess the functional form of the covariates in the intervention model (in the propensity score context) or the outcome model, (iii) adequately assess the covariate balance for so many variables. ******To address these challenges, there are four specific research objectives in this proposal. 1. To develop confounder selection approaches in a high dimensional setting incorporating the principles established in the causal inference literature. 2. To study the robustness of various data-adaptive methods in the context of model misspecification in a high dimensional setting. 3. To propose appropriate metrics for assessing the ‘covariate balance' in the context of propensity scores estimated from high-dimensional covariates. 4. To investigate the above issues when longitudinal data are available. These methods will be evaluated through theoretical developments, real-life applications, and via realistic simulations. ******I am positioned in a unique interdisciplinary research environment, as an Assistant Professor in the UBC Faculty of Medicine, a biostatistician at St. Paul's hospital, and an alumnus from the Statistics department, UBC, with close research ties with UBC and McGill. In this big-data era, there are huge demands for students with training in statistical modeling who can take causal structures into consideration while analyzing a large data set. Training of highly qualified personnel within an interdisciplinary environment is an essential component of this research. Trainees will receive training and access to high-quality research datasets and methodological and applied research questions that will have a real-life impact.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2022
  • 负责人:
    Karim, Mohammad
  • 依托单位:
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2021
  • 负责人:
    Karim, Mohammad
  • 依托单位:
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2020
  • 负责人:
    Karim, Mohammad
  • 依托单位:
Improving Causal Inference Methods in Statistics for Analyzing Big Data
  • 批准号:
    RGPIN-2018-05044
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.53万
  • 财政年份:
    2019
  • 负责人:
    Karim, Mohammad
  • 依托单位:
海外基金