Interpreting function of non-coding sequences with synthetic biology and machine learning
Interpreting function of non-coding sequences with synthetic biology and machine learning
批准号:
10065897
负责人:
Ryan Zachary Friedman
金额:
$3.15万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2023-06-30
关键词:
AddressAffectAlgorithm DesignArchitectureBase SequenceBiological AssayBiological ModelsBiologyCell Culture TechniquesCell LineCell modelCellsCellular AssayDNA SequenceDataData SetDiseaseEngineeringGene ExpressionGenesGenomeHumanImageIn VitroIndividualLibrariesMachine LearningMethodsModelingPerformancePhotoreceptorsPhysiologicalProductionRecommendationRegulationReporterRetinaStructureSystemTestingTissuesTrainingUntranslated RNAValidationVariantcell typecellular engineeringcomputer frameworkdesignexperienceexperimental studyfunctional genomicsgenetic variantgenomic datahigh throughput screeningimprovedin vivomachine learning algorithmpractical applicationprecision medicinesynthetic biologytranscription factor
中文摘要
项目总结/摘要
大多数疾病相关变异位于基因组的非编码区,并通过以下方式发挥影响:
对基因表达的影响然而,我们缺乏一个预测框架来解释这种非编码变体,
限制了基因组数据在精准医疗中的应用。我们也许可以用
新的机器学习算法,但到目前为止,机器学习在功能上的实际应用
基因组学由于两个主要挑战而受到限制。第一,训练数据集的规模和多样性
在功能基因组学中的应用比机器学习的应用小几个数量级,
成功的,如图像识别和产品推荐。第二个挑战是,如果训练数据
如果没有在适当的体外细胞模型中收集,则所得到的机器学习模型可能不
推广到相关的体内细胞类型。为了改进机器学习在非编码变体中的应用,
我建议同时解决训练数据集的有限大小和细胞培养模型的有效性。
机器学习的一个核心原则是模型性能随着数据的增加而提高。在目标1中,我建议
通过执行机器学习的迭代循环来增加训练数据的大小和多样性,
使用大规模平行报告基因测定(MPRAs)进行实验验证。我的方法的关键方面是
算法上设计每个连续的MPRA文库,以包含最有可能改善MPRA的序列。
下一轮的造型。我最近训练了我的第一个模型,这些数据是我从MPRA实验中收集的,
在哺乳动物光感受器中起作用的顺式调节序列。为了避免细胞系的任何问题,我
在离体发育的视网膜中进行这些实验,视网膜保留了适当的组织结构。
然而,与光感受器不同的是,大多数细胞类型在其天然生理学上在实验上是不可处理的。
上下文因此,重要的是确定体外细胞系如何再现体内顺式调节。在
目的2,我建议确定是否一个易处理的细胞培养模型可以概括的结果,从离体
视网膜我将使用现有的离体视网膜的MPRA数据作为标准,与
经工程改造以表达光感受器转录因子的组合的细胞系。我的目标是解决
将易处理的细胞系工程化以表达组织特异性转录因子可能是一种通用方法,
收集数据以训练机器学习模型,该模型推广到体内系统。成功完成
这些目标将产生一种通用的方法来增加功能基因组训练的规模和多样性
数据,并可能导致一种通用方法,用于生产实验上易于处理的机器学习系统
应用,最终帮助我们更好地将基因组数据应用于精准医疗。
英文摘要
PROJECT SUMMARY/ABSTRACT
Most disease-associated variants lie in non-coding regions of the genome and exert their influence through
effects on gene expression. However, we lack a predictive framework to interpret such non-coding variants,
limiting how genomic data is used in precision medicine. We may be able to interpret non-coding variants with
new machine learning algorithms, but so far the practical applications of machine learning in functional
genomics have been limited because of two major challenges. First, the size and diversity of training data sets
in functional genomics are orders of magnitude smaller than in applications where machine learning has been
successful, such as image recognition and product recommendation. A second challenge is that if training data
are not collected in an appropriate in vitro cellular model, then the resulting machine learning models may not
generalize to relevant in vivo cell types. To improve the application of machine learning to non-coding variants,
I propose to address both the limited size of training data sets and the efficacy of cell culture models.
A core principle of machine learning is that model performance improves with more data. In Aim 1, I propose to
increase the size and diversity of training data by performing iterative cycles of machine learning and
experimental validation with Massively Parallel Reporter Assays (MPRAs). The key aspect of my approach is to
algorithmically design each successive MPRA library to contain sequences that are most likely to improve the
next round of modeling. I recently trained my first model on data that I collected from MPRA experiments of
cis-regulatory sequences that function in mammalian photoreceptors. To avoid any issues with cell lines, I
performed these experiments in ex vivo developing retinas, which retain the appropriate tissue architecture.
However, unlike photoreceptors, most cell types are not experimentally tractable in their native physiological
context. Thus, it will be important to determine how well in vitro cell lines recapitulate in vivo cis-regulation. In
Aim 2, I propose to determine whether a tractable cell culture model can recapitulate results from ex vivo
retinas. I will use existing MPRA data from ex vivo retinas as a standard to compare against data collected in
cell lines engineered to express combinations of photoreceptor transcription factors. I aim to address whether
engineering tractable cell lines to express tissue-specific transcription factors might be a general approach for
collecting data to train machine learning models that generalize to in vivo systems. Successful completion of
these aims will produce a general approach to increase the size and diversity of functional genomic training
data, and may result in a general method for producing experimentally tractable systems for machine learning
applications, ultimately helping us better apply genomic data to precision medicine.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Interpreting function of non-coding sequences with synthetic biology and machine learning
-
批准号:10417177
-
项目类别:
-
资助金额:$3.08万
-
财政年份:2020
-
负责人:Ryan Zachary Friedman
-
依托单位:
Interpreting function of non-coding sequences with synthetic biology and machine learning
-
批准号:10177882
-
项目类别:
-
资助金额:$3.2万
-
财政年份:2020
-
负责人:Ryan Zachary Friedman
-
依托单位:
海外基金