课题基金 / 基金详情

Developing Machine Learning Models for the Analysis of Splicing Data in Large Heterogeneous Cohorts

Developing Machine Learning Models for the Analysis of Splicing Data in Large Heterogeneous Cohorts
开发机器学习模型来分析大型异构队列中的拼接数据
批准号:
10506326
负责人:
David Wang
金额:
$4.68万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2024-07-31

项目摘要

项目成果

David Wang的其他基金

相似基金

相关文献

中文摘要
翻译
摘要 从大型患者队列获得的RNA测序(RNASeq)数据分析可以揭示转录组学 与复杂疾病相关的扰动,并促进疾病亚型的鉴定。 这通常被框定为无监督学习任务,以发现RNASeq矩阵中的潜在结构。 基于基因表达或局部剪接变异(LSV)的定量。然而,有几个因素使 分析这种异构数据具有挑战性。首先,这样的数据集由在 可能采用不同测序方案和质量控制步骤的多个机构。这 在数据中引入混杂因素,如样本质量不一致或细胞类型比例可变 这会妨碍真实生物信号的检测。其次,在急性髓细胞白血病(AML)中, 剪接因子基因发生在一个子集的患者可能只会导致改变一个子集的共调节 剪接事件因此,代替基于所有转录组学测量样品之间的全局相似性, 特征,需要有效地识别由样本和拼接事件的子集定义的“瓦片”, 异常信号虽然已经提出了几种算法用于此任务,但它们未能克服许多缺点。 与建模拼接数据相关的计算挑战,并且不太适合处理缺失 价值观 为了通过减少假阳性发现和提高数据质量来促进异构拼接数据集的分析, 真正的生物信号,我们将首先开发一个模型来纠正RNA降解和细胞类型的影响, 混合物。然后,为了有效地鉴定以剪接事件为特征的AML亚型,并解释 拼接特定的建模挑战,我们提出棋盘(表征异质性的 通过搜索RNA数据集中的异常和异常值块的表达和剪接),一个非 用于瓦片的无监督发现的参数贝叶斯模型。我们将把我们的模型应用于合成数据集 并证明它优于几种基准方法。接下来,我们将展示它恢复的瓷砖,其特征在于 已知的和新的剪接畸变,其在多个AML患者队列中可重现。最后我们将 显示所发现的瓦片与药物对治疗的反应相关,指向翻译的 我们的调查结果的影响。
英文摘要
Abstract Analysis of RNA sequencing (RNASeq) data obtained from large patient cohorts can reveal transcriptomic perturbations that are associated with complex disease and facilitate the identification of disease subtypes. This is typically framed as an unsupervised learning task to discover latent structure in a matrix of RNASeq based quantification of gene expression or local splicing variations (LSVs). However, several factors make analysis of such heterogeneous data challenging. First, such datasets are comprised of samples processed at multiple institutions which might employ different sequencing protocols and quality control steps. This introduces confounding factors into the data like inconsistent sample quality or variable cell type proportions which can hinder detection of true biological signal. Second, in acute myeloid leukemia (AML), mutations in splice factor genes occurring in a subset of the patients may only result in alteration of a subset of coregulated splicing events. Thus, instead of measuring global similarity between samples based on all transcriptomic features, there is a need to efficiently identify “tiles”, defined by a subset of samples and splicing events with abnormal signals. Although several algorithms have been proposed for this task, they fail to overcome many of the computational challenges associated with modeling splicing data and are not well suited to handle missing values. To facilitate analysis of heterogeneous splicing datasets by reducing false positive discoveries and boosting true biological signal, we will first develop a model to correct for the effects of RNA degradation and cell type mixtures. Then in order to efficiently identify AML subtypes characterized by splicing events and account for splicing specific modeling challenges, we propose CHESSBOARD (Characterizing Heterogeneity of Expression and Splicing by Search for Blocks of Abnormalities and Outliers in RNA Datasets), a non- parametric Bayesian model for unsupervised discovery of tiles. We will apply our models to synthetic datasets and show it outperforms several baseline approaches. Next, we will show that it recovers tiles characterized by known and novel splicing aberrations which are reproducible in multiple AML patient cohorts. Finally, we will show that tiles discovered are correlated with drug response to therapeutics, pointing to the translational impact of our findings.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Developing Machine Learning Models for the Analysis of Splicing Data in Large Heterogeneous Cohorts
  • 批准号:
    10672974
  • 项目类别:
  • 资助金额:
    $3.46万
  • 财政年份:
    2021
  • 负责人:
    David Wang
  • 依托单位:
Developing Machine Learning Models for the Analysis of Splicing Data in Large Heterogeneous Cohorts
  • 批准号:
    10315802
  • 项目类别:
  • 资助金额:
    $4.6万
  • 财政年份:
    2021
  • 负责人:
    David Wang
  • 依托单位:
Neurodifferentiation/Stem Cell Unit
Neurodifferentiation/Stem Cell Unit
海外基金