Domain adaptation approaches to unify established and emerging sequencing technologies
Domain adaptation approaches to unify established and emerging sequencing technologies
批准号:
10643544
负责人:
Natalie Rose Davidson
金额:
$13.44万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-05-15 至 2025-04-30
关键词:
Acute Myelocytic LeukemiaAddressAdoptedAreaBiologicalBiological SciencesBiologyBone Marrow CellsCell modelCellsChromatinClinical SciencesColoradoComputational BiologyDataData SetEmerging TechnologiesExonsFaceFoundationsGene Expression ProfileHealthHigh-Throughput Nucleotide SequencingInformation TechnologyInstitutionInterdisciplinary StudyLearningMachine LearningMedicalMentorsMethodologyMethodsModelingPatternPhaseRNAResearchResearch PersonnelResourcesRheumatoid ArthritisSolidTechniquesTechnologyTissuesTrainingTranslational ResearchUniversitiesWorkbiological systemscell free DNAcell typeexperienceflexibilitygenetic signatureinterestmodel buildingnew technologynovel sequencing technologyprogenitorprogramsresponsetranscriptomicstranslational research programtreatment response
中文摘要
项目总结/摘要
测序技术的进步为从多个方面询问生物系统提供了新的机会。
视角然而,新技术的引入凸显了许多研究人员面临的一个问题:
缺少数据。技术和生物状态之间的观测缺失是一个经常观察到的问题
在计算生物学领域。这种缺失可能是由于技术的局限性,
或者是因为这项技术还没有被广泛采用。虽然一项技术可能
由于生物观测的高度稀疏性,有机会利用现有的互补数据,
一种用于估算缺失生物学观察结果的既定技术。
我们通过利用机器学习的新方法来解决这些问题,主要集中在
域适应技术。这些技术学习一个数据集中的模式,这些模式可以适应另一个数据集
数据集,实现跨技术信息共享。我们的提案提出了一个总体框架,
域自适应技术可用于将新兴技术与不同的技术结合起来。到
为了突出这种方法的广泛实用性,我们将这种模型应用于三个生物医学应用:1)预测
类风湿性关节炎中细胞类型特异性扰动反应; 2)从游离DNA预测组织来源
(cfDNA); 3)从急性髓性白血病(AML)中的无细胞DNA预测祖细胞特异性基因签名。
所提出的目标不仅统一了现有的和新兴的测序技术,
难以直接观察或不可行的新生物学。
这项研究建议建立在我的经验,使用统计方法的转录组数据。期间
K99第一阶段将需要我的指导团队在深度生成建模方面进行进一步的培训(凯西博士
格林)、单细胞数据建模(张凡博士)以及cfDNA和染色质可及性建模(张凡博士)。
Srinivas Ramachandran)。这项研究将在科罗拉多大学安舒茨医学中心进行
校园内,健康人工智能中心。在这个机构里,我可以进入科罗拉多临床中心,
转化科学研究所和RNA生物科学倡议,为建立一个
跨学科和翻译研究计划。通过这种培训和现有的机构资源,我
将有一个坚实的基础上,建立一个独立的研究计划,重点是领域适应
高通量测序技术的应用。
英文摘要
PROJECT SUMMARY / ABSTRACT
Advances in sequencing technologies provide new opportunities to interrogate biological systems from multiple
perspectives. However, the introduction of new technologies highlights a problem many researchers face:
missing data. Missing observations across technologies and biological states is a frequently observed problem
in the field of computational biology. This missingness can be a result of limitations in the technology, the rarity
of a biological state, or because the technology has not been widely adopted. While one technology may have
high sparsity in biological observations, there is an opportunity to leverage existing, complementary data from
an established technology to impute the missing biological observations.
We address these issues by utilizing new methodological advances in machine learning, primarily focusing on
domain adaptation techniques. These techniques learn patterns in one dataset that can be adapted to another
dataset, enabling cross-technology information sharing. Our proposal introduces a general framework in which
domain adaptation techniques can be used to unite an emerging technology with a different, but technology. To
highlight the broad utility of this approach, we apply this model to three biomedical applications: 1) Predict
cell-type-specific perturbation response in rheumatoid arthritis; 2) Predict tissue-of-origin from cell-free DNA
(cfDNA); 3) Predict progenitor-specific gene signatures from cell-free DNA in acute myeloid leukemia (AML).
The proposed aims not only unite existing and emerging sequencing technologies, but enable the discovery of
new biology that is difficult or infeasible to directly observe.
The research proposed builds on my experience in using statistical approaches for transcriptomic data. During
the K99 phase I will require further training from my mentoring team in deep generative modeling (Dr. Casey
Greene), modeling of single-cell data (Dr. Fan Zhang), and modeling of cfDNA and chromatin accessibility (Dr.
Srinivas Ramachandran). The research will be conducted at the University of Colorado, Anschutz Medical
Campus, in the Center for Health AI. In this institution, I will have access to the Colorado Clinical and
Translational Sciences Institute and the RNA Bioscience Initiative, which provide resources for building an
interdisciplinary and translational research program. With this training and available institutional resources, I
will have a solid foundation on which to build an independent research program focused on domain adaptation
applications for high-throughput sequencing technologies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金