Statistical learning and causal inference in high-dimensional genomics data across multiple information layers
Statistical learning and causal inference in high-dimensional genomics data across multiple information layers
批准号:
RGPIN-2022-03708
负责人:
Park, Yongjin
金额:
$1.38万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
越来越多的基因组数据被生成和共享,以研究人类生物学和复杂的疾病。由于数据的庞大数量和维度,统计机器学习方法已成为研究人员进行探索性数据分析并找到支持工作假设的证据的必然工具。本提案的主要目标是开发一种多组学机器学习(ML)方法来研究寻找人类特征的因果机制。与计算生物学中现有的专注于单一类型组学数据的ML方法不同,我们将强调在研究复杂表型时不同背景信息的重要性,以及在方法开发和分析中考虑多种数据模式的必要性。为了推进机器学习方法,特别是生物医学数据分析,我们寻求实现三个长期目标。(1)我们将提出一种新的数据集成和探索性建模方法,深入穿透多层次的生物信息流。我们将实现可扩展的贝叶斯推理方法,用于多模态单细胞数据集成和可解释的随机块模型,用于细胞-细胞,细胞-基因和基因-基因相互作用。(2)结合实际生成过程的知识,我们的机器学习方法将确定不同数据模式之间的因果机制,并最终邀请合作者在分子和细胞分辨率上剖析机制。对比学习方法将系统地结合多条科学/统计证据,以阐明“因果三角”中的因果机制,由此我们根据不同类型的对比,增加对某个感兴趣的假设的信心。(3)由于贝叶斯推理是包括我们在内的许多科学发现的关键计算步骤,我们将努力使推理方法广泛适用于多个科学领域。值得注意的是,我们将实现一个黑箱学习算法,该算法同时采用个人层面和汇总统计数据。使用我们的机器学习方法,与生物医学研究小组合作,我们将提出人类生物学中的基本问题:人类特征的自然分布是什么?我们能描述表型变异的主轴吗?不同的特征是如何相互联系的?从健康状态转变为病理状态的关键因素是什么?在实现统计学的长期目标的同时,我的小组将寻求分析大量的现实世界数据,为众多的科学问题提供定量的答案。我们将根据科学兴趣将七个HQPs分成三个工作组:癌症生物学(2个HQPs),单细胞方法学(3个HQPs),免疫紊乱组(2个HQPs)。我们与英属哥伦比亚大学、维多利亚大学、耶鲁大学、麻省理工学院等世界一流的实验实验室合作。
英文摘要
An increasingly large amount of genomics data are generated and shared to study human biology and complex disorders. Due to the sheer volume and dimensions of data, statistical machine learning methods have become an inevitable tool for researchers to conduct exploratory data analysis and find evidence to support working hypotheses. The main objective of this proposal is to develop a multi-omics machine learning (ML) approach to the study of finding causal mechanisms of human traits. Unlike existing ML methods in computational biology focusing on a single type of omics data, we will emphasize the importance of diverse contextual information in studying complex phenotypes and the necessity to consider multiple data modalities in method development and analysis. To advance ML methods, specializing in biomedical data analysis, we seek to achieve three long-term objectives. (1) We will present a new approach for data integration and exploratory modelling, penetrating deeply through multiple layers of biological information flows. We will implement scalable Bayesian inference methods for multi-modal single-cell data integration and interpretable stochastic block models for cell-cell, cell-gene, and gene-gene interactions. (2) Incorporating the knowledge of the actual generative process, our ML methods will ascertain causal mechanisms across different data modalities and eventually invite collaborators to dissect the mechanisms at a molecular and cellular resolution. A contrastive learning approach will systematically combine multiple lines of scientific/statistical evidence to elucidate causal mechanisms in "causal triangulation," whereby we increase confidence for a certain hypothesis of interest in light of different types of contrasts. (3) Since Bayesian inference is a crucial computational step to many scientific discoveries, including ours, we will strive to make inference methods widely applicable to multiple scientific domains. Notably, we will implement a black-box learning algorithm that takes both individual-level and summary statistics data. Using our ML approach, in collaboration with biomedical research groups, we will ask fundamental questions in human biology: What are natural distributions of human traits? Can we characterize principal axes of phenotypic variation? How are different traits linked with one another? What are the key contributors that make the transitions from healthy to pathological states? While achieving the long-term goals in statistics, my group will seek to analyze a massive amount of real-world data to give quantitative answers to numerous scientific questions. We will organize a total of seven HQPs into three working groups based on scientific interests: cancer biology (2 HQPs), single-cell methodology (3 HQPs), immune disorder groups (2 HQPs). We collaborate with world-class experimental laboratories in University of British Columbia, University of Victoria, Yale, and MIT.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical learning and causal inference in high-dimensional genomics data across multiple information layers
-
批准号:DGECR-2022-00445
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2022
-
负责人:Park, Yongjin
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
儿童音乐能力发展对语言与社会认知能力及脑发育的影响
-
批准号:31971003
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:南云
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
多场景网络学习中基于行为-情感-主题联合建模的学习者兴趣挖掘关键技术研究
-
批准号:61702207
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2017
-
负责人:刘智
-
依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
-
批准号:61672236
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2016
-
负责人:王骏
-
依托单位: