课题基金 / 基金详情

Multivariate Modelling and Inference of Dependent and High-Dimensional Data in Recent Genetic Studies

Multivariate Modelling and Inference of Dependent and High-Dimensional Data in Recent Genetic Studies
最近遗传学研究中相关和高维数据的多变量建模和推理
批准号:
RGPIN-2019-06727
负责人:
Oualkacha, Karim
金额:
$1.46万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Oualkacha, Karim的其他基金

相似基金

相关文献

中文摘要
翻译
统计遗传学正在经历向大数据的过渡,这与应用统计学的几个分支正在经历的情况相同,而且随着高通量基因组实验的到来,这种过渡正在加速。多年来,我的研究活动相应地发生了变化,以应对新兴基因组技术带来的新挑战。这项研究计划包括三个应对此类挑战的创新研究主题,重点是开发依赖基因组数据的多变量统计工具。下一代测序(NGS)技术现在提供了DNA序列变异的详尽目录,挑战变成了理解这些变异的表型后果。第一个目标是关于NGS关联研究中多种表型的最佳使用。多个相关的表型通常测量相同的潜在特征,并且可以与疾病诊断具有更直接的关系。通过Copula模型提供对表型依赖结构的灵活建模,并利用核技巧(即探索表型-基因型关系的机器学习方法),本研究轴展示了相关表型和遗传变异之间潜在关系的广泛建模。该框架将增加识别导致人类复杂疾病的新基因变异的能力,这可能有助于更好地理解疾病病因。数据正规化对于高通量基因组学数据很有吸引力,以便检测/选择结果的相关预测因子的较小子集。这样的数据也表现出异质性,这是许多研究人员感兴趣的,但现有的预测模型往往忽略了这一点。第二个目标是在受罚的稳健回归模型中实施现代计算统计学的支柱算法,以捕获受试者内部的相关性,并为相关基因组数据选择/检测相关的异质预测因子。这样的预测模型将是一个有用的工具,可以建立遗传风险评分,这对风险分层和临床决策非常有用。第三个目标是一个长期目标,将来自第一和第二研究轴的统计工具结合起来,以建立一个统一的基于Copula的关联框架,能够识别不同的遗传变异,同时提供表型依赖的灵活建模。它将更深入地了解遗传变异是如何解释表型变异的;这被称为“遗漏遗传力”问题,大多数现有的遗传学研究都遇到了这一问题。缺乏有效的统计方法来分析现代基因组数据是基因组研究界更好地了解相关生物学所面临的一大瓶颈。我坚信,我在这项提案中提出的策略将非常有助于分析和整合这些复杂的数据,并将有助于最大限度地发挥其效用。
英文摘要
Statistical genetics is undergoing the same transition to big data that several branches of applied statistics are experiencing, and this transition is accelerating with the advent of high-throughput genomic experiments. My research activities have evolved accordingly throughout the years to address the new challenges brought on by the emerging genomic technologies. This research program consists of three innovative research themes dealing with such challenges, with a focus on the development of multivariate statistical tools for dependent genomics data. Next-generation sequencing (NGS) technology is now providing an exhaustive catalog of DNA sequence variation and the challenge becomes understanding the phenotypic consequences of these variants. The first objective deals with the optimal use of multiple phenotypes in NGS association studies. Multiple correlated phenotypes often measure the same underlying trait and can bear a more direct relationship with the disease diagnosis. By providing flexible modeling of the phenotypes dependence structure via copula models, and exploiting Kernel trick (i.e. machine learning methods for exploring phenotype-genotype relationship), this research axis lays out a broad modeling of the underlying relationship between correlated phenotypes and genetic variants. The framework will increase power in identifying novel genetic variants responsible for human complex diseases, which may help to better understand disease etiology. Data-regularization is appealing for high-throughput genomics data to detect/select a smaller subset of relevant predictors for an outcome. Such data display also heterogeneity which is of interest to many researchers but it tends to be overlooked by existing predictive models. The second objective implements pillar algorithms of modern computational statistics within penalized robust regression models to capture within-subject dependence and select/detect relevant heterogeneous predictors for dependent genomics data. Such prediction models will be a useful tool to build genetic risk scores that can be very useful for risk stratification and clinical decision-making. The third objective is a long-term goal which couples statistical tools from the first and second research axes to build a unified copula-based association framework capable to identify heterogeneous genetic variants while providing flexible modeling of the phenotypes dependence. It will gain more insight on how genetic variation is explaining the phenotypic variation; this is known as "missing heritability" problem and is encountered by most existing genetic studies. The lack of efficient statistical methods to analyze modern genomics data is a major bottleneck faced by the genomics research community to better understand the related biology. I strongly believe that the strategies I propose in this proposal will be very useful for analyzing and integrating such complex data, and will help with maximizing their utility.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multivariate Modelling and Inference of Dependent and High-Dimensional Data in Recent Genetic Studies
  • 批准号:
    RGPIN-2019-06727
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2021
  • 负责人:
    Oualkacha, Karim
  • 依托单位:
Multivariate Modelling and Inference of Dependent and High-Dimensional Data in Recent Genetic Studies
  • 批准号:
    RGPIN-2019-06727
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2020
  • 负责人:
    Oualkacha, Karim
  • 依托单位:
Multivariate Modelling and Inference of Dependent and High-Dimensional Data in Recent Genetic Studies
  • 批准号:
    RGPIN-2019-06727
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2019
  • 负责人:
    Oualkacha, Karim
  • 依托单位:
Multivariate Modelling and Inference in Genetic Studies
  • 批准号:
    433266-2013
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2018
  • 负责人:
    Oualkacha, Karim
  • 依托单位:
国内基金
海外基金
Improving modelling of compact binary evolution.
  • 批准号:
    10903001
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2009
  • 负责人:
    史蒂芬
  • 依托单位: