DINOS: Data INspired Oligo Synthesis for DNA Data Storage

DINOS: Data INspired Oligo Synthesis for DNA Data Storage
复制标题

DOI:
10.1145/3510853
复制
发表时间:
2022-03
期刊:
ACM Journal on Emerging Technologies in Computing Systems (JETC)
影响因子:
--
通讯作者:
Kevin Volkel;Kyle J Tomek;Albert J. Keung;James M. Tuck
Kevin Volkel;Kyle J Tomek;Albert J. Keung;James M. Tuck
中科院分区:
其他
文献类型:
--
作者:
Kevin Volkel;Kyle J Tomek;Albert J. Keung;James M. Tuck

文献摘要

相似文献

随着人们对基于 DNA 的信息存储的兴趣不断增长,合成成本已被确定为一个关键瓶颈。一个潜在的方向是调整数据综合。数据链往往由一小组重复出现的码字序列组成,并且它们包含较长的重复数据序列。为了利用这些特性,我们提出了一个名为 DINOS 的新框架。 DINOS 由三个关键部分组成:(i)第一个是分层链组装算法,受基因组装技术启发,可以从一小组原始块组装任意数据链。 (ii) 汇编算法依赖于我们关于如何构造基元块的新颖公式,跨越一组码字和悬垂的各种有用配置。每个基元块都是一个码字,两侧有一对突出部分,这些突出部分是由循环配对过程创建的,该过程使基元块的数量保持较小。理论上,使用这些原始块,可以组装任意长度的任何数据链。我们展示了一个最小的二进制代码系统,只有六个基本块,并且我们概括了我们的过程以支持任意一组悬垂和代码字。 (iii)我们利用分层组装方法来识别冗余序列并合并产生它们的反应,以使组装更加高效。我们评估 DINOS 并描述其主要特征。例如,可以通过增加突出端的数量或码字的数量来减少形成链所需的反应数量,但增加突出端的数量相对于增加码字来说具有较小的优势,同时需要显着更少的基元块。然而,通过增加码字的数量可以更多地提高密度。我们还发现,即使组装的最小数据片段为 16 位,简单的冗余合并技术也能够将解压缩和压缩数据的反应平均分别减少 90.6% 和 41.2%。通过发现更多冗余的简单填充启发式方法,我们可以进一步将相同操作点的反应减少,对于解压缩和压缩数据,平均分别高达 91.1% 和 59%。与之前的通用基因组装技术相比,我们的方法可提供高达 80% 的密度。最后,在对合成成本的分析中,我们使用从头合成制作 1 GB 卷,与仅使用从头合成制作原始块并使用 DINOS 进行组装相比,我们估计 DINOS 比从头合成便宜 105 倍。
As interest in DNA-based information storage grows, the costs of synthesis have been identified as a key bottleneck. A potential direction is to tune synthesis for data. Data strands tend to be composed of a small set of recurring code word sequences, and they contain longer sequences of repeated data. To exploit these properties, we propose a new framework called DINOS. DINOS consists of three key parts: (i) The first is a hierarchical strand assembly algorithm, inspired by gene assembly techniques that can assemble arbitrary data strands from a small set of primitive blocks. (ii) The assembly algorithm relies on our novel formulation for how to construct primitive blocks, spanning a variety of useful configurations from a set of code words and overhangs. Each primitive block is a code word flanked by a pair of overhangs that are created by a cyclic pairing process that keeps the number of primitive blocks small. Using these primitive blocks, any data strand of arbitrary length can be assembled, theoretically. We show a minimal system for a binary code with as few as six primitive blocks, and we generalize our processes to support an arbitrary set of overhangs and code words. (iii) We exploit our hierarchical assembly approach to identify redundant sequences and coalesce the reactions that create them to make assembly more efficient. We evaluate DINOS and describe its key characteristics. For example, the number of reactions needed to make a strand can be reduced by increasing the number of overhangs or the number of code words, but increasing the number of overhangs offers a small advantage over increasing code words while requiring substantially fewer primitive blocks. However, density is improved more by increasing the number of code words. We also find that a simple redundancy coalescing technique is able to reduce reactions by 90.6% and 41.2% on average for decompressed and compressed data, respectively, even when the smallest data fragments being assembled are 16 bits. With a simple padding heuristic that finds even more redundancy, we can further decrease reactions for the same operating point up to 91.1% and 59% for decompressed and compressed data, respectively, on average. Our approach offers greater density by up to 80% over a prior general purpose gene assembly technique. Finally, in an analysis of synthesis costs in which we make 1 GB volume using de novo synthesis versus making only the primitive blocks with de novo synthesis and otherwise assembling using DINOS, we estimate DINOS as 105× cheaper than de novo synthesis.