课题基金 / 基金详情

CRII: III: Computational framework for disparate data integration to study cancer drug resistance

CRII: III: Computational framework for disparate data integration to study cancer drug resistance
CRII:III:用于研究癌症耐药性的不同数据整合的计算框架
批准号:
1850360
负责人:
Sha Cao
金额:
$17.47万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-08-15 至 2022-12-31

项目摘要

项目成果

Sha Cao的其他基金

相似基金

相关文献

中文摘要
翻译
随着高通量生物技术的出现,人们现在可以使用不同类型的研究对象进行多种生物检测来研究生物系统。多个生物检测、研究对象和暴露条件的卷积产生了丰富的关于生物系统的信息,同时也给如何整合异质数据源和提取仅从单个数据集无法获得的足够知识提出了巨大的挑战。根据卷积的复杂程度,当对一组相似的研究对象梳理来自多个生物测定的信息时,通常使用直接数据集成,而当涉及到结合相对不同的数据类型的能力时,通过知识转移的间接集成通常是有吸引力的。直接信息集成依赖于一致且注释良好的数据结构,并且数据集之间的连接通常很容易理解。由于所合并的数据的大小、格式和维度不同,这些方法必须应对许多计算挑战。在生物学研究中,研究人员倾向于使用不同的模型来研究他们的系统,这些模型根据物种(人与鼠)、成分(组织与细胞系、单细胞)和/或暴露条件的不同而不同。因此,生成的数据集来自不同的特征空间和/或不同的样本分布,其中直接集成变得不可行。这些高度不同的数据集都有可能为正在进行的研究问题提供关键的补充信息,因此迫切需要构建一种针对高通量生物测试数据的迁移学习方法,以便在未来的研究中重新利用现有的异质数据集。为了解决这些挑战,该项目将开发新的计算方法类,用于直接信息集成和间接知识转移,并最终利用各种生物测试数据集之间的结构和关系来更好地了解生物系统,例如癌症耐药机制。研究组将通过实施以下两个目标来实现他们的目标。首先,他们将开发一种新的配方,用于对多组学数据进行信息蒸馏,以便知识可以很容易地移植到未来的研究中,否则,由于技术和平台偏见,这些研究将被阻止。该框架包括用于定性表示相干签名的有监督稀疏聚类方法和用于检测局部低等级结构的联合聚类方法。其次,他们将开发一种新的迁移学习方法,将从源域学习的结构规则应用到任何目标域。关键的假设是,结构规则不受技术和平台偏差的影响,因此它们是知识转移的理想载体。该项目预计将开发新的计算工具,能够有效地探索各种不同的数据集,并通过最大限度地利用现有信息并充分利用从中获得的知识来证实我们对生物/生物医学系统的理解,从而最大限度地减少重新收集新训练数据的成本。因此,该项目将产生深远的经济和社会影响。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
With the advent of high-throughput biotechnology, people can now investigate a biological system with multiple bioassays using diverse types of study objects. The convolution of multiple bioassays, study objects and exposure conditions produces a wealth of rich information about the biological system, and at the same time, poses great challenge on how to integrate heterogeneous data sources and extract sufficient knowledge that cannot be gained from any single dataset alone. Depending on the complexity level of the convolution, direct data integration is often used when combing information from multiple bioassays on a cohort of similar study objects, and indirect integration via knowledge transfer is generally attractive when it comes to the power of incorporating relatively more diverse data types. Direct information integration relies on concordant and well annotated data structures, and the connections between datasets are usually easily understood. These methodologies have to meet many computational challenges owing to different sizes, formats and dimensionalities of the data being integrated. In biological studies, researchers tend to study their system using models varying in terms of species (human vs. mouse), compositions (tissue vs. cell lines, single cells) and/or exposures conditions. Thus, the datasets generated are drawn from different feature space and/or different sample distributions, where direct integration becomes infeasible. These highly disparate datasets may each have the potential to provide complimentary information key to the research question being carried out, and thus it is also urgent to construct a transfer learning method tailored for high-throughput bioassay data so that existing heterogeneous datasets could be re-purposed in a future study.To address these challenges, this project will develop new classes of computational methods for direct information integration and indirect knowledge transfer, and ultimately leverage structures and relations among various bioassay datasets for better understanding of a biological system, for example, cancer drug resistance mechanism. The research team will achieve their goals through exerting the following two objectives. First, they will develop a novel formulation for information distillation on multiple -omics data so that knowledge could be easily transplantable to future studies, which would otherwise be prevented due to technical and platform bias. The framework consists of a supervised sparse clustering method for qualitative representation of coherent signatures, together with a co-clustering approach to detect local low rank structures. Second, they will develop a novel transfer learning method by imposing the structural regularities learnt from the source domains to any target domain. The key assumption is that the structural regularities are invariant to technical and platform bias, so they are ideal vehicle for knowledge transfer. The project is expected to develop novel computational tools that can effectively explore a wide range of heterogeneous datasets, and it has great potential to minimize the cost on recollecting new training data, by maximizing utilization of existing information and fully using the knowledge derived therefrom to substantiate our understanding of a biological/biomedical system. Hence the project will have far-reaching economic and societal impacts.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/bibm49941.2020.9313483
发表时间: 2020-12
期刊: 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
影响因子: --
作者: [Changlin Wan;D. Jia;Yue Zhao;Wennan Chang;Sha Cao;Xiao Wang;Chi Zhang]
通讯作者: Changlin Wan;D. Jia;Yue Zhao;Wennan Chang;Sha Cao;Xiao Wang;Chi Zhang
DOI: --
发表时间: 2022-08
期刊: Proceedings of machine learning research
影响因子: --
作者: [Changlin Wan;Pengtao Dang;Tong Zhao;Y. Zang;Chi Zhang-;Sha Cao]
通讯作者: Changlin Wan;Pengtao Dang;Tong Zhao;Y. Zang;Chi Zhang-;Sha Cao
DOI: --
发表时间: 2020-07
期刊: ArXiv
影响因子: --
作者: [Changlin Wan;Wennan Chang;Tong Zhao;Sha Cao;Chi Zhang]
通讯作者: Changlin Wan;Wennan Chang;Tong Zhao;Sha Cao;Chi Zhang
Response to ‘Letter to the Editor: on the stability and internal consistency of component-wise sparse mixture regression based clustering’, Zhang et al.
回应“致编辑的信:关于基于组件的稀疏混合回归聚类的稳定性和内部一致性”,Zhang 等人。
DOI: 10.1093/bib/bbac262
发表时间: 2022
期刊: Briefings in Bioinformatics
影响因子: 9.5
作者: [Chang, Wennan, Zhang, Chi, Cao, Sha]
通讯作者: Cao, Sha
共 8 条
    CAREER: Disentangled learning of high dimensional biomedical data in the presence of inherent heterogeneity
    • 批准号:
      2145314
    • 项目类别:
      Standard Grant
    • 资助金额:
      $59.65万
    • 财政年份:
      2022
    • 负责人:
      Sha Cao
    • 依托单位:
    国内基金
    海外基金
    基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
    基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
    • 批准号:
      JCZRLH202600780
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2026
    • 负责人:
    • 依托单位:
    白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
    • 批准号:
      2026JJ82690
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2026
    • 负责人:
      张卓
    • 依托单位:
    基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
    • 批准号:
      2026JJ30130
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2026
    • 负责人:
      张二军
    • 依托单位: