DNA representations and generalization performance of sequence-to-expression models

DNA representations and generalization performance of sequence-to-expression models
复制标题

DOI:
10.1101/2024.02.06.579067
复制
发表时间:
2024-02
期刊:
bioRxiv
影响因子:
--
通讯作者:
Yuxin Shen;Grzegorz Kudla;D. Oyarzún
Yuxin Shen;Grzegorz Kudla;D. Oyarzún
中科院分区:
其他
文献类型:
--
作者:
Yuxin Shen;Grzegorz Kudla;D. Oyarzún

文献摘要

相似文献

对生物制品日益增长的需求促使许多人努力设计细胞,以最大产量生产异种蛋白。大规模并行报告分析的最新进展可以提供适合训练机器学习模型的数据,并支持具有优化蛋白质表达表型的微生物菌株的设计。表现最好的序列到表达模型已经在单热编码上进行了训练,这是核苷酸序列的一种机制不可知的表示。然而,尽管这些模型具有出色的局部预测能力,但它们在远离训练数据的地方泛化预测的能力有限。在这里,我们展示了遗传构造库可以根据所选择的序列表示具有本质上不同的聚类结构,并证明可以利用这种差异来提高泛化性能。使用来自大肠杆菌的大型序列到表达式数据集,我们表明非深度回归器和在单热编码上训练的卷积神经网络无法泛化预测,并且使用最先进的大型语言模型学习表征也难以达到域外精度。相比之下,我们表明,尽管它们的局部性能较差,但机制序列特征(如密码子偏倚、核苷酸含量或mRNA稳定性)在模型泛化方面提供了有希望的收益。我们探索了几种将不同特征集集成到单个预测模型中的策略,包括特征叠加、集成模型叠加和几何叠加,几何叠加是一种基于图卷积神经网络的新架构。我们的工作表明,领域不可知和领域感知序列特征的整合为提高序列到表达模型的质量提供了一条尚未探索的途径,并促进了它们在生物技术和制药领域的应用。
The increasing demand for biological products drives many efforts to engineer cells that produce heterologous proteins at maximal yield. Recent advances in massively parallel reporter assays can deliver data suitable for training machine learning models and sup-port the design of microbial strains with optimized protein expression phenotypes. The best performing sequence- to-expression models have been trained on one-hot encodings, a mechanism-agnostic representation of nucleotide sequences. Despite their excellent local pre-dictive power, however, such models suffer from a limited ability to generalize predictions far away from the training data. Here, we show that libraries of genetic constructs can have substantially different cluster structure depending on the chosen sequence representation, and demonstrate that such differences can be leveraged to improve generalization perfor-mance. Using a large sequence- to-expression dataset from Escherichia coli, we show that non-deep regressors and convolutional neural networks trained on one-hot encodings fail to generalize predictions, and that learned representations using state-of-the-art large language models also struggle with out-of-domain accuracy. In contrast, we show that despite their poorer local performance, mechanistic sequence features such as codon bias, nucleotide con-tent or mRNA stability, provide promising gains on model generalization. We explore several strategies to integrate different feature sets into a single predictive model, including feature stacking, ensemble model stacking, and geometric stacking, a novel architecture based on graph convolutional neural networks. Our work suggests that integration of domain-agnostic and domain-aware sequence features offers an unexplored route for improving the quality of sequence- to-expression models and facilitate their adoption in the biotechnology and phar-maceutical sectors.