Multidimensional data organization and random access in large-scale DNA storage systems

Multidimensional data organization and random access in large-scale DNA storage systems
复制标题

大规模DNA存储系统中的多维数据组织和随机访问

DOI:
10.1016/j.tcs.2021.09.021
复制
发表时间:
2021
影响因子:
1.1
通讯作者:
Reif, John
Reif, John
中科院分区:
计算机科学4区
文献类型:
--
作者:
Song, Xin;Shah, Shalin;Reif, John

文献摘要

相似文献

DNA具有令人印象深刻的物理密度和分子级编码能力,是构建持久数据存档存储系统的有前途的基底。为了从DNA存储中检索数据,最近的实现通常依赖于精心设计的正交PCR引物的大型文库,这从根本上限制了实际DNA存储的容量和可扩展性。这项工作结合了嵌套式和半嵌套式PCR,使多维数据组织和随机访问大型DNA存储。我们的策略有效地推动了DNA存储容量的极限,并大大减少了高效PCR随机访问所需的正交引物的数量。我们的设计仅使用k n引物来唯一地处理n k个数据编码寡核苷酸。该体系结构固有地支持各种定义明确的PCR随机访问模式,这些模式可以被定制为以简单的类似数据库的格式(诸如行、列、表和数据条目块)来组织和保存底层DNA编码的数据结构和关系。我们设计了四维DNA存储的计算机PCR实验,以说明16种不同的随机访问模式的机制,每个模式需要不超过两个PCR反应来选择性地扩增各种大小的目标数据集。为了更好地近似物理系统,我们制定了基于经验分布的数学模型来分析移液,PCR偏差和PCR随机性对大型DNA存储多维数据查询性能的影响。
With impressive physical density and molecular-scale coding capacity, DNA is a promising substrate for building long-lasting data archival storage systems. To retrieve data from DNA storage, recent implementations typically rely on large libraries of meticulously designed orthogonal PCR primers, which fundamentally limit the capacity and scalability of practical DNA storage. This work combines nested and semi-nested PCR to enable multidimensional data organization and random access in large DNA storage. Our strategy effectively pushes the limit of DNA storage capacity and dramatically reduces the number of orthogonal primers needed for efficient PCR random access. Our design uses only k⁎ n primers to uniquely address n k data-encoding oligos. The architecture inherently supports various well-defined PCR random-access patterns that can be tailored to organize and preserve the underlying DNA-encoded data structures and relations in simple database-like formats such as rows, columns, tables, and blocks of data entries. We design in silico PCR experiments of a four-dimensional DNA storage to illustrate the mechanisms of sixteen different random-access patterns each requiring no more than two PCR reactions to selectively amplify a target dataset of various sizes. To better approximate the physical system, we formulate mathematical models based on empirical distributions to analyze the effect of pipetting, PCR bias, and PCR stochasticity on the performance of multidimensional data queries from large DNA storage.