DNA Structure Design Is Improved Using an Artificially Expanded Alphabet of Base Pairs Including Loop and Mismatch Thermodynamic Parameters.

DNA Structure Design Is Improved Using an Artificially Expanded Alphabet of Base Pairs Including Loop and Mismatch Thermodynamic Parameters.
复制标题

使用人工扩展的碱基对字母表(包括环和错配热力学参数)改进 DNA 结构设计。

DOI:
10.1101/2023.06.06.543917
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Mathews,DavidH
Mathews,DavidH
中科院分区:
--
文献类型:
--
作者:
Pham,TuanM;Miffin,Terrel;Sun,Hongying;Sharp,KennethK;Wang,Xiaoyu;Zhu,Mingyi;Hoshika,Shuichi;Peterson,RaymondJ;Benner,StevenA;Kahn,JasonD;Mathews,DavidH

文献摘要

相似文献

我们发现,通过将碱基配对字母表从A-T和G-C扩展到包括2-氨基-8-(1′-β-d-2′-脱氧呋喃核糖基)-咪唑并[1,2-a]-1,3,5-三嗪-(8H)-4-酮和6-氨基-3-(1′-β-d-2′-脱氧呋喃核糖基)-5-硝基-(1H)-吡啶-2-酮(简称PandZ)之间的配对,改进了DNA二级结构的计算机设计。为了获得在设计中包含P-Z对所需的热力学参数,我们进行了47次光学熔化实验,并将结果与以前的工作相结合,以拟合P-Z对和G-Z摆动对的自由能和焓最近邻折叠参数。我们发现G-Z对具有与A-T对相当的稳定性,因此应该作为碱基对包括在结构预测和设计算法中。此外,我们外推了环、末端错配和悬挂末端参数的集合以包括P和Z核苷酸。这些参数被整合到RNAstructure软件包中用于二级结构预测和分析。使用RNA结构设计程序,我们解决了Eterna使用ACGT字母表或补充P-Z对提出的100个设计问题中的99个。正如标准化系综缺陷(NED)所评估的那样,扩展字母表降低了序列折叠成脱靶结构的倾向。在99个提供Eterna播放器解决方案的案例中,有91个案例的NED值相对于Eterna示例解决方案的NED值有所改善。含P-Z的设计的平均NED值为0.040,显著低于仅含标准DNA的设计的0.074,并且包含P-Z对减少了收敛于设计所需的时间。这项工作提供了一个样本管道,用于将任何扩展的字母核苷酸纳入预测和设计工作流程。
We show thatin silicodesign of DNA secondary structures is improved by extending the base pairing alphabet beyond A–T and G–C to include the pair between 2-amino-8-(1′-β-d-2′-deoxyribofuranosyl)-imidazo-[1,2-a]-1,3,5-triazin-(8H)-4-one and 6-amino-3-(1′-β-d-2′-deoxyribofuranosyl)-5-nitro-(1H)-pyridin-2-one, abbreviated asPandZ. To obtain the thermodynamic parameters needed to include P–Z pairs in the designs, we performed 47 optical melting experiments and combined the results with previous work to fit free energy and enthalpy nearest neighbor folding parameters for P–Z pairs and G–Z wobble pairs. We find G–Z pairs have stability comparable to that of A–T pairs and should therefore be included as base pairs in structure prediction and design algorithms. Additionally, we extrapolated the set of loop, terminal mismatch, and dangling end parameters to include the P and Z nucleotides. These parameters were incorporated into the RNAstructure software package for secondary structure prediction and analysis. Using the RNAstructure Design program, we solved 99 of the 100 design problems posed by Eterna using the ACGT alphabet or supplementing it with P–Z pairs. Extending the alphabet reduced the propensity of sequences to fold into off-target structures, as evaluated by the normalized ensemble defect (NED). The NED values were improved relative to those from the Eterna example solutions in 91 of 99 cases in which Eterna-player solutions were provided. P–Z-containing designs had average NED values of 0.040, significantly below the 0.074 of standard-DNA-only designs, and inclusion of the P–Z pairs decreased the time needed to converge on a design. This work provides a sample pipeline for inclusion of any expanded alphabet nucleotides into prediction and design workflows.