Applying machine learning to predict viral assembly for adeno-associated virus capsid libraries.

Applying machine learning to predict viral assembly for adeno-associated virus capsid libraries.
复制标题

应用机器学习预测腺相关病毒衣壳文库的病毒组装。

DOI:
10.1016/j.omtm.2020.11.017
复制
发表时间:
2021-03-12
期刊:
Molecular therapy. Methods & clinical development
影响因子:
--
通讯作者:
Zolotukhin S
Zolotukhin S
中科院分区:
其他
文献类型:
--
作者:
Marques AD;Kummer M;Kondratov O;Banerjee A;Moskalenko O;Zolotukhin S

文献摘要

参考文献

被引文献

相似文献

机器学习(ML)可以帮助病毒基因治疗领域的新发现。具体而言,通过复杂衣壳文库的下一代测序(NGS)收集的大数据是数据分析和预测中失去潜力的一个特别突出的来源。此外,基于腺相关病毒(AAV)的衣壳文库作为选择用于基因治疗载体的候选物的工具正变得越来越受欢迎。这些更高复杂性的AAV衣壳文库先前已经在体内创建和选择;然而,使用ML计算机算法的计算机模拟分析可以增加更智能和更稳健的文库用于选择。在这项研究中,在病毒组装之前和之后收集的AAV衣壳文库的数据用于训练ML算法。我们发现,两种机器学习计算机算法,人工神经网络(ANN)和支持向量机(SVM),可以被训练来预测未知的衣壳变体是否可以组装成可行的病毒样结构。使用构建的最精确的模型,模拟文库构建中的假设突变模式以表明N495、G546和I554在AAV 2衍生的衣壳中的重要性。最后,使用ML衍生的数据生成两个比较文库,以生物学验证这些发现并证明ML在载体设计中的预测能力。Marques等人开发一个生物信息学管道,实施机器学习来预测基于AAV 2的氨基酸序列是否会产生可行的衣壳。进行这些预测的能力可能有助于研究人员通过选择具有较少死端变体的库来优化AAV基因治疗载体文库。
Machine learning (ML) can aid in novel discoveries in the field of viral gene therapy. Specifically, big data gathered through next-generation sequencing (NGS) of complex capsid libraries is an especially prominent source of lost potential in data analysis and prediction. Furthermore, adeno-associated virus (AAV)-based capsid libraries are becoming increasingly popular as a tool to select candidates for gene therapy vectors. These higher complexity AAV capsid libraries have previously been created and selected in vivo; however, in silico analysis using ML computer algorithms may augment smarter and more robust libraries for selection. In this study, data of AAV capsid libraries gathered before and after viral assembly are used to train ML algorithms. We found that two ML computer algorithms, artificial neural networks (ANNs), and support vector machines (SVMs), can be trained to predict whether unknown capsid variants may assemble into viable virus-like structures. Using the most accurate models constructed, hypothetical mutation patterns in library construction were simulated to suggest the importance of N495, G546, and I554 in AAV2-derived capsids. Finally, two comparative libraries were generated using ML-derived data to biologically validate these findings and demonstrate the predictive power of ML in vector design. Marques et al. develop a bioinformatics pipeline implementing machine learning to predict whether AAV2-based amino acid sequences will produce viable capsids. The ability to make these predictions may facilitate researchers to optimize AAV gene therapy vector libraries by selecting pools with fewer dead-end variants in silico.
DOI: 10.1016/j.cell.2012.04.012
发表时间: 2012-06-22
期刊: Cell
影响因子: 64.5
作者:
Hopf TA;Colwell LJ;Sheridan R;Rost B;Sander C;Marks DS
通讯作者: Marks DS
DOI: 10.1093/bioinformatics/bty166
发表时间: 2018-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Khurana, Sameer;Rawi, Reda;Mall, Raghvendra
通讯作者: Mall, Raghvendra
DOI: 10.1128/jvi.80.2.821-834.2006
发表时间: 2006-01-01
影响因子: 5.4
作者:
Lochrie, MA;Tatsuno, GP;Colosi, P
通讯作者: Colosi, P
DOI: 10.1093/bioinformatics/btx350
发表时间: 2017-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Jimenez, J.;Doerr, S.;De Fabritiis, G.
通讯作者: De Fabritiis, G.
DOI: 10.1038/mt.2008.100
发表时间: 2008-07-01
期刊: MOLECULAR THERAPY
影响因子: 12.4
作者:
Li, Wuping;Asokan, Aravind;Samulski, Richard J.
通讯作者: Samulski, Richard J.