Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
批准号:
RGPIN-2018-05147
负责人:
Zhang, Qingrun
金额:
$2.26万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
了解表型变化的遗传基础和基于基因型预测表型是遗传学领域的长期目标。今天,可扩展的工具,即,允许使用个人计算机中有限的存储器快速分析非常大的数据的软件是基因组大数据时代的新兴需求。我的研究计划的长期目标是开发新的统计模型及其可扩展的实现,以促进基因型-表型映射和预测。背景高通量测序技术的最新进展,包括全外显子组测序,RNA-Seq(用于转录组)和亚硫酸氢盐-Seq(用于甲基化组),在遗传学相关领域产生了令人兴奋的结果。然而,缺乏允许多尺度组学数据集无缝集成的工具,用于在生物学相关和有意义的背景下精确预测表型。特别是,基因-基因相互作用还没有得到充分的表征和利用的预测。此外,在即将到来的大数据时代,缺乏可扩展的工具,允许对大型数据集进行有效的统计分析,而不需要具有非常大内存的机器。基于我以前的工作,确定基因-基因相互作用和实现可扩展的软件,我将专注于三个短期目标:(1)通过整合多尺度组学数据,利用基因型-表型数据识别基因相互作用;贝叶斯网络和频繁项目集挖掘将被集成来实现这一目标。(2)建立一个多基因表型预测,整合基因相互作用;组LASSO的扩展将实现这一点。(3)使用基于磁盘的解决方案创建可扩展的软件来实现上述统计模型,即,内存虚拟化和计算机科学中的大页面技术。它将在磁盘上存储大量数据,同时允许快速计算,就像数据驻留在主存中一样。将使用来自1,001拟南芥基因组计划的多尺度组学数据和在阿尔伯塔产生的多尺度组学植物数据。 冲击这项研究不仅将为统计遗传学提供一个新的理论框架,而且还将提供新的计算工具,以帮助实验科学家进行基因作图项目。此外,它将使农业和卫生从业人员能够改善对表型的预测,使加拿大的食品生产和卫生系统受益。实际上,我的软件代表了一个可扩展的解决方案,用于大数据分析,使用最少的内存。这将特别关系到许多加拿大研究小组,他们不能立即获得高性能计算设施。该计划将培养HQP进行生物信息学和生物统计学分析,以充分利用未来的基因组大数据。
英文摘要
Understanding the genetic basis of phenotypic changes and predicting phenotypes based on genotypes are long-standing goals of the field of genetics. Today, scalable tools, i.e., software that allows rapid analysis of very large data using limited memory in a personal computer, is an emerging requirement in the era of genomic big-data. The long-term goal of my research program is to develop novel statistical models and their scalable implementations to facilitate genotype-phenotype mappings and predictions. Background. Recent advances in high-throughput sequencing technologies including whole-exome sequencing, RNA-Seq (for transcriptome) and Bisulfite-Seq (for methylome) have created an excitement in genetics-related areas. However, there is a lack of tools that allow the seamless integration of multi-scale omics datasets for the precise prediction of phenotypes in a biologically relevant and meaningful context. In particular, gene-gene interactions have not been fully characterized and utilized in predictors. Moreover, in the coming big-data era, there is a lack of scalable tools that permit the effective statistical analysis of large datasets without requiring a machine with very large memory.Objectives. Building upon my previous work of identifying gene-gene interactions and implementing scalable software, I will focus on three short-term objectives: (1) Identify gene interactions using genotype-phenotype data by integrating multi-scale omics data; Bayesian Network and Frequent Itemset Mining will be integrated to achieve this goal. (2) Build a polygenic phenotype predictor that integrates gene interactions; an extension of Group LASSO will be implemented for this. (3) Create scalable software implementing the aforementioned statistical models using disk-based solutions, i.e., memory virtualization and huge-page techniques in computer science. It will store large data on disk while allowing rapid calculation as if the data resided in main memory. The multi-scale omics data from the 1,001 Arabidopsis Genomes Project and the multi-scale omics plant data generated in Alberta will be used. Impact. The proposed research will not only provide a new theoretical framework of statistical genetics, but will also furnish novel computational tools to assist experimental scientists carrying out gene mapping projects. Moreover, it will enable practitioners in agriculture and health to improve predictions of phenotype, benefiting Canadian food productions and the health system. Practically, my software represents a scalable solution for big-data analyses with minimum memory usage. This will be particularly relevant to many Canadian research groups that do not have immediate access to high-performance computing facilities. This program will train HQP to carry out bioinformatics and biostatistics analyses to fully utilize the future genomic big-data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2021
-
负责人:Zhang, Qingrun
-
依托单位:
A GPU Server for Integration of Machine Learning in Mathematics and Statistics Research and Training
-
批准号:RTI-2021-00675
-
项目类别:Research Tools and Instruments
-
资助金额:$10.92万
-
财政年份:2020
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2020
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2019
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:RGPIN-2018-05147
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.26万
-
财政年份:2018
-
负责人:Zhang, Qingrun
-
依托单位:
Statistical models and computational tools for gene-gene interaction analyses by utilizing multi-scale omics
-
批准号:DGECR-2018-00061
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2018
-
负责人:Zhang, Qingrun
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
河北南部地区灰霾的来源和形成机制研究
-
批准号:41105105
-
项目类别:青年科学基金项目
-
资助金额:25.0万元
-
批准年份:2011
-
负责人:王丽涛
-
依托单位:
保险风险模型、投资组合及相关课题研究
-
批准号:10971157
-
项目类别:面上项目
-
资助金额:24.0万元
-
批准年份:2009
-
负责人:胡亦钧
-
依托单位:
RKTG对ERK信号通路的调控和肿瘤生成的影响
-
批准号:30830037
-
项目类别:重点项目
-
资助金额:190.0万元
-
批准年份:2008
-
负责人:陈雁
-
依托单位:
新型手性NAD(P)H Models合成及生化模拟
-
批准号:20472090
-
项目类别:面上项目
-
资助金额:23.0万元
-
批准年份:2004
-
负责人:王乃兴
-
依托单位: