课题基金 / 基金详情

CAREER: Disentangled learning of high dimensional biomedical data in the presence of inherent heterogeneity

CAREER: Disentangled learning of high dimensional biomedical data in the presence of inherent heterogeneity
职业:在存在固有异质性的情况下对高维生物医学数据进行解缠学习
批准号:
2145314
负责人:
Sha Cao
金额:
$59.65万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2027-07-31

项目摘要

项目成果

Sha Cao的其他基金

相似基金

相关文献

中文摘要
翻译
高通量技术的进步正在彻底改变生物医学研究领域,产生了大规模的患者来源的分子数据集,包括匹配的多组学数据、单细胞或空间分辨组学数据以及纵向组学数据。尽管分析这种复杂数据以进一步理解人类疾病状态的分子基础的趋势日益增长,但固有的生物异质性正在挑战将给定疾病系统作为统一实体治疗的传统范式:患者可能形成具有可变临床结果的不同疾病亚型,并且从同一患者样本收集的细胞可能显示不同的表型状态。面对高维分子特征的生物学和临床可解释性仍然是另一个挑战。拟议的项目将解决以下关键挑战:1)同时检测具有生物学/临床意义的患者/细胞亚群,并提取其独特的定义基因组特征; 2)从多种模式的数据中进行综合学习,具有较强的系统批量效应,或随着时间的推移收集。因此,所提出的工作具有很高的潜力,发现新的知识,并诱导新的方法显着的健康数据科学研究。该项目将产生广泛适用于整个数据科学的算法,如高维约简,聚类和数据集成,以满足广泛的研究和行业需求。该项目提出了解开学习的新想法,以同时解决三种最流行的生物医学数据类型的高维和固有的异构性问题:多组学、组织分辨组学和纵向组学数据。将解决以下三个关键挑战:1)通过新颖的监督聚类方法利用多组学数据来检测生物学/临床上有意义的患者亚组; 2)通过基于泊松的低维嵌入模型发现噪声组织分辨组学数据中的异质细胞群体和多个组织样本的多任务学习;以及3)通过融合学习模型使用纵向和高维组学数据来表征异质疾病轨迹。对于所有这三种情况,提出了一个共同的主题,解开学习,以确保生物/临床分析结果的可解释性。受试者内部和跨受试者的固有异质性被划分为不同的亚组或亚群,每个亚组或亚群的特征在于从高维特征空间内提取的其独特特征。所有提出的方法都在学术界和工业界具有广泛的实用性。在教育方面,拟议的同伴学习教育模块,基于R Shiny的工具开发研究,以及针对高中和本科生的数据科学夏季研讨会,不仅可以作为激励和留住学生的平台,还可以培训他们未来的STEM劳动力角色。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The advancement in high throughput technologies is revolutionizing the fields of biomedical research, giving rise to large-scale patient-derived molecular datasets, including matched multi-omics data, single cell or spatially resolved -omics data, as well as longitudinal -omics data. Despite the growing trend of analyzing such complex data to further our understanding of the molecular basis in states of human disease, the inherent biological heterogeneity is challenging the conventional paradigm to treat a given diseased system as a uniform entity: patients may form different disease subtypes present with variable clinical outcomes, and cells collected from the same patient sample may display different phenotypic states. Biological and clinical interpretability in the face of the high dimensional molecular features remain to be another challenge. The proposed project will address key challenges in: 1) simultaneously detecting biologically/clinically meaningful patient/cell subgroups and extracting their unique defining genomic features; and 2) integrated learning from data of multiple modalities, with strong systematic batch effect, or collected over time. Thus, the proposed work has high potential to discover new knowledge and to induce new approaches significant to health data science research. The project will result in algorithms, such as high dimensionality reduction, clustering and data integration, that are broadly applicable across the whole of data science, to address a wide range of research and industry needs.This project proposes novel ideas of disentangled learning to simultaneously address the high dimensionality and inherent heterogeneity issues for three most popular biomedical data types: multi-omics, tissue resolved -omics, and longitudinal -omics data. The following three critical challenges are to be addressed: 1) Detecting biologically/clinically meaningful patient subgroups by leveraging multi-omics data through novel supervised clustering methods; 2) Discovering heterogeneous cell populations in noisy tissue resolved omics data and multi-task learning of multiple tissue samples through a Poisson based low dimensional embedding model; and 3) Characterizing heterogeneous disease trajectories using longitudinal and high dimensional -omics data through a fusion learning model. For all three scenarios, a common theme of disentangled learning is proposed to ensure the biological/clinical interpretability of the analysis results. The inherent heterogeneity within and across the subjects are divided to distinct subgroups or subpopulations, each of which is characterized by their unique features extracted from within a high dimensional feature space. All proposed methods are embraced with extensive utility in both academia and industry. Educationally, the proposed peer learning education module, R Shiny based tool development research, and summer workshops on data science, all targeting high school and undergraduate students, can well serve as a platform not only for inspiring and retaining students, but also training them for future STEM workforce roles. Thus, the proposed project is promising to have far-reaching educational impacts.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2022-08
期刊: Proceedings of machine learning research
影响因子: --
作者: [Changlin Wan;Pengtao Dang;Tong Zhao;Y. Zang;Chi Zhang-;Sha Cao]
通讯作者: Changlin Wan;Pengtao Dang;Tong Zhao;Y. Zang;Chi Zhang-;Sha Cao
Rethink Boolean matrix factorization by bias-aware disentangled representation learning
通过偏差感知解缠表示学习重新思考布尔矩阵分解
DOI: --
发表时间: 2023
期刊: KDD
影响因子: --
作者: [Dang, P, Zhao T, Salama P, Wang Y, Zhang C]
通讯作者: Zhang C
CRII: III: Computational framework for disparate data integration to study cancer drug resistance
  • 批准号:
    1850360
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.47万
  • 财政年份:
    2019
  • 负责人:
    Sha Cao
  • 依托单位:
海外基金