Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
批准号:
RGPIN-2018-05147
负责人:
Zhang, Qingrun
金额:
$2.26万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
了解表型变化的遗传基础并根据基因类型预测表型是遗传学领域长期追求的目标。今天,可扩展的工具,即允许使用个人计算机中有限的内存快速分析非常大的数据的软件,是基因组大数据时代的一种新兴需求。我的研究计划的长期目标是开发新的统计模型及其可扩展的实现,以促进基因-表型映射和预测。背景资料。高通量测序技术的最新进展,包括全外显子组测序、RNA-Seq(转录组)和Bisulfite-Seq(甲基组),在遗传学相关领域引起了轰动。然而,缺乏工具允许无缝集成多尺度组学数据集,以便在生物学相关和有意义的背景下精确预测表型。特别是,基因-基因相互作用还没有被充分描述和用于预测。此外,在即将到来的大数据时代,缺乏可扩展的工具来允许对大数据集进行有效的统计分析,而不需要具有非常大内存的机器。在我之前识别基因-基因交互作用和实施可扩展软件的工作的基础上,我将专注于三个短期目标:(1)通过整合多尺度组学数据,使用基因-表型数据识别基因交互作用;贝叶斯网络和频繁项集挖掘将整合以实现这一目标。(2)构建一个整合了基因互作的多基因表型预测因子;为此将实现群体套索的扩展。(3)使用基于磁盘的解决方案,即内存虚拟化和计算机科学中的巨页技术,创建可扩展的软件,实现上述统计模型。它将大量数据存储在磁盘上,同时允许快速计算,就像数据驻留在主内存中一样。将使用来自1,001个拟南芥基因组项目的多尺度组学数据和在艾伯塔省产生的多尺度组学植物数据。冲击力。这项拟议的研究不仅将提供统计遗传学的新理论框架,还将提供新的计算工具来帮助实验科学家进行基因图谱项目。此外,它将使农业和卫生从业者能够提高对表型的预测,使加拿大的食品生产和卫生系统受益。实际上,我的软件代表了一种可扩展的大数据分析解决方案,内存使用量最小。这对于许多不能立即使用高性能计算设施的加拿大研究小组来说尤其重要。该计划将培训HQP进行生物信息学和生物统计分析,以充分利用未来的基因组大数据。
英文摘要
Understanding the genetic basis of phenotypic changes and predicting phenotypes based on genotypes are long-standing goals of the field of genetics. Today, scalable tools, i.e., software that allows rapid analysis of very large data using limited memory in a personal computer, is an emerging requirement in the era of genomic big-data. The long-term goal of my research program is to develop novel statistical models and their scalable implementations to facilitate genotype-phenotype mappings and predictions. Background. Recent advances in high-throughput sequencing technologies including whole-exome sequencing, RNA-Seq (for transcriptome) and Bisulfite-Seq (for methylome) have created an excitement in genetics-related areas. However, there is a lack of tools that allow the seamless integration of multi-scale omics datasets for the precise prediction of phenotypes in a biologically relevant and meaningful context. In particular, gene-gene interactions have not been fully characterized and utilized in predictors. Moreover, in the coming big-data era, there is a lack of scalable tools that permit the effective statistical analysis of large datasets without requiring a machine with very large memory.Objectives. Building upon my previous work of identifying gene-gene interactions and implementing scalable software, I will focus on three short-term objectives: (1) Identify gene interactions using genotype-phenotype data by integrating multi-scale omics data; Bayesian Network and Frequent Itemset Mining will be integrated to achieve this goal. (2) Build a polygenic phenotype predictor that integrates gene interactions; an extension of Group LASSO will be implemented for this. (3) Create scalable software implementing the aforementioned statistical models using disk-based solutions, i.e., memory virtualization and huge-page techniques in computer science. It will store large data on disk while allowing rapid calculation as if the data resided in main memory. The multi-scale omics data from the 1,001 Arabidopsis Genomes Project and the multi-scale omics plant data generated in Alberta will be used. Impact. The proposed research will not only provide a new theoretical framework of statistical genetics, but will also furnish novel computational tools to assist experimental scientists carrying out gene mapping projects. Moreover, it will enable practitioners in agriculture and health to improve predictions of phenotype, benefiting Canadian food productions and the health system. Practically, my software represents a scalable solution for big-data analyses with minimum memory usage. This will be particularly relevant to many Canadian research groups that do not have immediate access to high-performance computing facilities. This program will train HQP to carry out bioinformatics and biostatistics analyses to fully utilize the future genomic big-data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2021
-
负责人:Zhang, Qingrun
-
依托单位:
A GPU Server for Integration of Machine Learning in Mathematics and Statistics Research and Training
-
批准号:RTI-2021-00675
-
项目类别:Research Tools and Instruments
-
资助金额:$10.92万
-
财政年份:2020
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2020
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2019
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2018
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:DGECR-2018-00061
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2018
-
负责人:Zhang, Qingrun
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
河北南部地区灰霾的来源和形成机制研究
-
批准号:41105105
-
项目类别:青年科学基金项目
-
资助金额:25.0万元
-
批准年份:2011
-
负责人:王丽涛
-
依托单位:
保险风险模型、投资组合及相关课题研究
-
批准号:10971157
-
项目类别:面上项目
-
资助金额:24.0万元
-
批准年份:2009
-
负责人:胡亦钧
-
依托单位:
RKTG对ERK信号通路的调控和肿瘤生成的影响
-
批准号:30830037
-
项目类别:重点项目
-
资助金额:190.0万元
-
批准年份:2008
-
负责人:陈雁
-
依托单位:
新型手性NAD(P)H Models合成及生化模拟
-
批准号:20472090
-
项目类别:面上项目
-
资助金额:23.0万元
-
批准年份:2004
-
负责人:王乃兴
-
依托单位: