Hold out the genome: a roadmap to solving the cis-regulatory code

Hold out the genome: a roadmap to solving the cis-regulatory code
复制标题

DOI:
10.1038/s41586-023-06661-w
复制
发表时间:
2023-12-13
期刊:
影响因子:
64.8
通讯作者:
Taipale,Jussi
Taipale,Jussi
中科院分区:
综合性期刊1区
文献类型:
--
作者:
de Boer,Carl G.;Taipale,Jussi

文献摘要

相似文献

基因表达受转录因子调节,这些转录因子共同作用以读取顺式调节DNA序列。“顺式调节密码”-细胞如何解释DNA序列以确定何时,何地以及多少基因应该表达-已被证明是极其复杂的。最近,功能基因组学测定和机器学习的规模和分辨率的进步使得破译这一代码取得了实质性进展。然而,如果模型只在基因组序列上训练,那么这种调控代码可能永远不会得到解决;同源区域很容易导致对预测性能的高估,而我们的基因组太短,序列多样性不足,无法学习所有相关参数。幸运的是,随机合成的DNA序列可以测试比我们基因组中存在的更大的序列空间,设计的DNA序列可以进行有针对性的查询,以最大限度地改善模型。由于相同的生物化学原理被用于解释DNA,无论其来源如何,在这些合成数据上训练的模型可以预测基因组活性,通常比基因组训练的模型更好。在这里,我们提供了该领域的前景,并提出了一个路线图,通过机器学习和使用合成DNA的大规模并行检测相结合来解决这些监管代码。
Gene expression is regulated by transcription factors that work together to readcis-regulatory DNA sequences. The ‘cis-regulatory code’ — how cells interpret DNA sequences to determine when, where and how much genes should be expressed — has proven to be exceedingly complex. Recently, advances in the scale and resolution of functional genomics assays and machine learning have enabled substantial progress towards deciphering this code. However, thecis-regulatory code will probably never be solved if models are trained only on genomic sequences; regions of homology can easily lead to overestimation of predictive performance, and our genome is too short and has insufficient sequence diversity to learn all relevant parameters. Fortunately, randomly synthesized DNA sequences enable testing a far larger sequence space than exists in our genomes, and designed DNA sequences enable targeted queries to maximally improve the models. As the same biochemical principles are used to interpret DNA regardless of its source, models trained on these synthetic data can predict genomic activity, often better than genome-trained models. Here we provide an outlook on the field, and propose a roadmap towards solving thecis-regulatory code by a combination of machine learning and massively parallel assays using synthetic DNA.