III: Small: Collaborative Research: Combinatorial Collaborative Clustering for Simultaneous Patient Stratification and Biomarker Identification
III: Small: Collaborative Research: Combinatorial Collaborative Clustering for Simultaneous Patient Stratification and Biomarker Identification
批准号:
1812699
负责人:
Mingyuan Zhou
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-15 至 2023-08-31
中文摘要
现代高通量测序(HTS)技术产生了丰富的高维生物医学数据。在研究复杂的、动态的、随机的、异质的生命和疾病系统时,样本的维度(特征数量)通常比HTS数据中的样本数量高得多。这些HTS数据在带来巨大的统计和计算挑战的同时,也为合作研究带来了独特的机会,将其转化为临床精确医学。这个项目将开发新的贝叶斯方法和计算工具,用于组合协作聚类,目标是两个基本的生物医学应用:肿瘤分层和预测性生物标志物识别。与现有的肿瘤分层和生物标记物识别黑盒算法相比,贝叶斯组合协同聚类框架能够同时对特定肿瘤亚型进行肿瘤分层和生物标记物识别,从而从机理上理解复杂疾病的异质性。捕捉到的分子图谱和疾病亚型之间的相互关系可能为疾病细胞机制提供深入的见解,并有可能开发个性化的疾病预后和治疗策略。该项目的跨学科性质,以及计划的课程开发和推广活动,将为本科生和研究生提供极好的培训机会,使他们掌握生物医学研究的量化技能,拥有前所未有的大量生物医学数据。该项目的核心是一个新的贝叶斯统计框架的理论和计算基础,以将现有的大规模公开生物医学数据集,如TCGA(癌症基因组图谱)和ICGC(国际癌症基因组联盟),转化为精确(个性化)的疾病诊断和预后。基于现代HTS数据,将开发一类新的二进制和计数数据分析模型,用于组合协作聚类(CCC),以实现可重复性和准确的肿瘤分层和生物标志物识别。这里“组合”意味着每个集群将被定义在一个特征子集上,该特征子集将通过新颖的组合分析从所有可能的特征组合中选择;和“协作”意味着每个集群是由其集群成员如何表达其所选特征子集来协作定义的。首先,CCC不是定义集群中心和距离度量来根据所有特征对患者进行分层,而是同时将特定于集群的特征识别为在执行患者分层时显示相似轮廓模式的生物标记物。因此,通过在从数万个特征中选择的一小部分特征上计算患者群下的样本的预测似然,它减轻了“维度的诅咒”,并显著提高了重复性。其次,它还通过将各种类型的数据与潜在计数联系起来,实现了混合类型HTS数据的自然集成。最后,所提出的基于计数建模的推理算法只对非零元素进行计算,因此产生了非常高效的稀疏矩阵分析方法,通常是在高温超导数据中。除了理论和计算上的优点,CCC还提供了一个灵活的概率计算框架来识别和表征肿瘤亚型或亚克隆,从而导致更有效的个性化预后和治疗设计。建议的CCC方法将首先在TCGA和ICGC数据上进行评估,然后应用于与首席研究员正在进行的癌症和免疫疾病研究的生物医学合作者的合作研究中。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern high-throughput sequencing (HTS) technologies produce rich high-dimensional biomedical data. When studying complex, dynamic, stochastic, and heterogeneous life and disease systems, the dimensionality (number of features) of samples is typically much higher than the number of samples in HTS data. Such HTS data, while imposing significant statistical and computational challenges, bring unique opportunities for collaborative research to translate them to clinical precision medicine. This project will develop novel Bayesian methods and computational tools for combinatorial collaborative clustering targeting at two fundamental biomedical applications: tumor stratification and predictive biomarker identification. Compared to existing black-box algorithms for tumor stratification and biomarker identification, the proposed Bayesian combinatorial collaborative clustering framework enables simultaneous tumor stratification and biomarker identification for specific tumor subtypes, so that mechanistic understanding of heterogeneity of complex diseases can be obtained. The captured interrelationships between molecular profile patterns and disease subtypes may provide deep insights into disease cellular mechanisms and have the potential of developing personalized disease prognosis and therapeutic strategies. The interdisciplinary nature of this project, together with the planned curriculum development and outreach activities, will provide excellent training opportunities for both undergraduate and graduate students, preparing them with the quantitative skills in biomedical research with unprecedented big biomedical data.The core of this project is the theoretic and computational foundation of a novel Bayesian statistical framework to translate existing large-scale publicly available biomedical datasets, such as TCGA (The Cancer Genome Atlas) and ICGC (International Cancer Genome Consortium), to precision (personalized) disease diagnosis and prognosis. A new class of binary and count data analysis models will be developed for Combinatorial Collaborative Clustering (CCC) based on modern HTS data to achieve reproducible and accurate tumor stratification and biomarker identification. Here "combinatorial' means that each cluster will be defined over a subset of features, which will be selected from all possible feature combinations, via novel combinatorial analysis; and "collaborative" means that each cluster is collaboratively defined by how its cluster members express their selected subset of features. First, rather than defining cluster centers and a distance metric to stratify patients based on all features, CCC simultaneously identifies cluster-specific features as biomarkers that show similar profile patterns when performing patient stratification. Hence, with the predictive likelihood of a sample under a patient cluster calculated over a small subset of features selected from tens of thousands of them, it alleviates "the curse of dimensionality" and substantially improves reproducibility. Second, it also enables natural integration of mixed-type HTS data by linking various types of data to latent counts. Finally, the proposed count modeling based inference algorithms only compute for non-zero elements and therefore lead to extremely efficient analytic methods for sparse matrices, often the case in HTS data. In addition to the theoretic and computational merit, CCC provides a flexible probabilistic computational framework to identify and characterize tumor subtypes or subclones, which leads to more effective personalized prognosis and therapeutic design. The proposed CCC methods will be first evaluated on the TCGA and ICGC data, and then be applied to the collaborative research with the principal investigator's ongoing biomedical collaborators on cancer and immunological disease studies.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(47)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2021-05
期刊:
影响因子:
--
作者:
[A. Dimitriev;Mingyuan Zhou]
通讯作者:
A. Dimitriev;Mingyuan Zhou
Dadaneh, Siamak Zamani, et al. "Pairwise supervised hashing with Bernoulli variational auto-encoder and self-control gradient estimator
达达内 (Dadaneh)、西亚马克·扎马尼 (Siamak Zamani) 等人。
DOI:
--
发表时间:
2020
期刊:
Conference on Uncertainty in Artificial Intelligence
影响因子:
--
作者:
[Dadaneh, S.Z.]
通讯作者:
Dadaneh, S.Z.
DOI:
--
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
作者:
[Mohammadreza Armandpour;Mingyuan Zhou]
通讯作者:
Mohammadreza Armandpour;Mingyuan Zhou
DOI:
--
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
作者:
[He Zhao;Piyush Rai;Lan Du;Wray L. Buntine;Mingyuan Zhou]
通讯作者:
He Zhao;Piyush Rai;Lan Du;Wray L. Buntine;Mingyuan Zhou
ALLSH: Active Learning Guided by Local Sensitivity and Hardness
ALLSH:以局部敏感性和硬度为指导的主动学习
DOI:
10.18653/v1/2022.findings-naacl.99
发表时间:
2022
期刊:
2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics
影响因子:
--
作者:
[Zhang, S., Gong, C., Liu, X., He, P., Chen, W., Zhou, M.]
通讯作者:
Zhou, M.
共 46 条
Collaborative Research: III: Medium: Conditional Transport: Theory, Methods, Computation, and Applications
-
批准号:2212418
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2022
-
负责人:Mingyuan Zhou
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: