CAREER: Mining biological functions from single cell multi-omics data
CAREER: Mining biological functions from single cell multi-omics data
批准号:
2047631
负责人:
Chi Zhang
金额:
$79.85万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-04-15 至 2026-03-31
中文摘要
Biological functional activities include intracellular functions such as transcriptional regulation, metabolism, and signaling transduction, and intercellular activities such as cell-cell interactions. With the advent of single cell multi-omics (scMulti-seq) biotechnology, researchers can study the biological functions of a complex biological system at the cellular resolution. The integrative analysis of scMulti-seq data and multiple study objects produces a wealth of rich information that enables the characterization of species or tissue specific biological functions, and at the same time, poses great challenge on how to identify and extract biologically meaningful data patterns. Though substantial amount of efforts has been made to interpret data patterns in single cell multi omics data, most of the existing methods focused on unsupervised learning in a completely data driven manner without considering the rich existing knowledge. In addition, depending on the types of biological functions, their underlying mathematical representation forms are different in scMulti-seq data. This calls for systems biology models and machine learning concepts to target true biological functions from scMulti-seq data. The first challenge to study biological functions from scMulti-seq data is to derive the data patterns that correspond to true biological functions and develop proper computational models for specific biological mechanisms and pathways. The second challenge lies in the difficulty of knowledge representation and sharing across the studies for different species, tissue types and experimental conditions. There remains an urgent need to integrate knowledge derived from disparate data sources to optimize the biological functional modeling, such that the learned knowledge could be utilized to study other biological systems or data types and promote the generation of new hypotheses. The PI’s long-term career goal is to develop mathematical formulations and computational methods to model biological functions from multi-omics data. This project will develop new mathematical models and an advanced computational framework to optimize the mining of biological functions, by integrating scMulti-seq data with context specific and general knowledge derived from independent data sets or experiments. The PI's research team will achieve the goals through the following three objectives. First, a novel subspace representation model will be developed to identify transcriptional regulation and functional gene modules. The proposed method will be empowered by a novel local low-rank matrix detection method to detect gene co regulation modules and a meta-learning framework to optimize results interpretation. Second, the PI's research team will develop a new graph neural network architecture to estimate cell-wise functional activities for flux carrying networks and a graph data clustering method to identify cell groups with varied functional states and distinct pathways. Thirdly, a knowledge graph will be constructed to represent the biological functions derived from scMulti-seq data, which enables the integration of independent knowledge derived from literature data and development of new biological hypotheses. The project is expected to deliver novel computational tools that can effectively explore biological functions from a wide range of heterogeneous datasets, and it could provide new capabilities for functional interpretation of individual data sets by maximizing the utilization of existing scMulti-seq and literature data, and reasoning of new biological hypotheses and mechanisms. Educationally, the scientific discoveries, including developed methods and biological knowledge, will be seamlessly integrated into an online educational knowledge base for large-scale public engagement, and will also lead to new project-based interdisciplinary training for high school, undergraduate and graduate students. The results of this project can be found at: https://zcslab.github.io/.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
英文摘要
Biological functional activities include intracellular functions such as transcriptional regulation, metabolism, and signaling transduction, and intercellular activities such as cell-cell interactions. With the advent of single cell multi-omics (scMulti-seq) biotechnology, researchers can study the biological functions of a complex biological system at the cellular resolution. The integrative analysis of scMulti-seq data and multiple study objects produces a wealth of rich information that enables the characterization of species or tissue specific biological functions, and at the same time, poses great challenge on how to identify and extract biologically meaningful data patterns. Though substantial amount of efforts has been made to interpret data patterns in single cell multi omics data, most of the existing methods focused on unsupervised learning in a completely data driven manner without considering the rich existing knowledge. In addition, depending on the types of biological functions, their underlying mathematical representation forms are different in scMulti-seq data. This calls for systems biology models and machine learning concepts to target true biological functions from scMulti-seq data. The first challenge to study biological functions from scMulti-seq data is to derive the data patterns that correspond to true biological functions and develop proper computational models for specific biological mechanisms and pathways. The second challenge lies in the difficulty of knowledge representation and sharing across the studies for different species, tissue types and experimental conditions. There remains an urgent need to integrate knowledge derived from disparate data sources to optimize the biological functional modeling, such that the learned knowledge could be utilized to study other biological systems or data types and promote the generation of new hypotheses. The PI’s long-term career goal is to develop mathematical formulations and computational methods to model biological functions from multi-omics data. This project will develop new mathematical models and an advanced computational framework to optimize the mining of biological functions, by integrating scMulti-seq data with context specific and general knowledge derived from independent data sets or experiments. The PI's research team will achieve the goals through the following three objectives. First, a novel subspace representation model will be developed to identify transcriptional regulation and functional gene modules. The proposed method will be empowered by a novel local low-rank matrix detection method to detect gene co regulation modules and a meta-learning framework to optimize results interpretation. Second, the PI's research team will develop a new graph neural network architecture to estimate cell-wise functional activities for flux carrying networks and a graph data clustering method to identify cell groups with varied functional states and distinct pathways. Thirdly, a knowledge graph will be constructed to represent the biological functions derived from scMulti-seq data, which enables the integration of independent knowledge derived from literature data and development of new biological hypotheses. The project is expected to deliver novel computational tools that can effectively explore biological functions from a wide range of heterogeneous datasets, and it could provide new capabilities for functional interpretation of individual data sets by maximizing the utilization of existing scMulti-seq and literature data, and reasoning of new biological hypotheses and mechanisms. Educationally, the scientific discoveries, including developed methods and biological knowledge, will be seamlessly integrated into an online educational knowledge base for large-scale public engagement, and will also lead to new project-based interdisciplinary training for high school, undergraduate and graduate students. The results of this project can be found at: https://zcslab.github.io/.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1186/s13046-021-02046-x
发表时间:
2021-08-10
期刊:
Journal of experimental & clinical cancer research : CR
影响因子:
--
作者:
[Gampala S, Shah F, Lu X, Moon HR, Babb O, Umesh Ganesh N, Sandusky G, Hulsey E, Armstrong L, Mosely AL, Han B, Ivan M, Yeh JJ, Kelley MR, Zhang C, Fishel ML]
通讯作者:
Fishel ML
DOI:
10.1109/icdm51629.2021.00013
发表时间:
2021-09
期刊:
2021 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
作者:
[Wennan Chang;Pengtao Dang;Changlin Wan;Xiaoyu Lu;Yue Fang;Tong Zhao;Y. Zang;Bo Li;Chi Zhang;Sha Cao]
通讯作者:
Wennan Chang;Pengtao Dang;Changlin Wan;Xiaoyu Lu;Yue Fang;Tong Zhao;Y. Zang;Bo Li;Chi Zhang;Sha Cao
DOI:
10.1101/2022.03.04.482927
发表时间:
2022
期刊:
Genomics proteomics and bioinformatics
影响因子:
9.5
作者:
[Zhou, Yi, Chang, Wennan, Lu, Xiaoyu, Wang, Jin, Zhang, Chi, Xu, Ying]
通讯作者:
Xu, Ying
国内基金
海外基金
基于Genome mining技术研究抑制表皮葡萄球菌生物膜形成的次级代谢产物
-
批准号:21242003
-
项目类别:专项基金项目
-
资助金额:10.0万元
-
批准年份:2012
-
负责人:昌军
-
依托单位: