Removing the Bottleneck: Introducing cMatch - A Lightweight Tool for Construct-Matching in Synthetic Biology.

Removing the Bottleneck: Introducing cMatch - A Lightweight Tool for Construct-Matching in Synthetic Biology.
复制标题

DOI:
10.3389/fbioe.2021.785131
复制
发表时间:
2021
影响因子:
5.7
通讯作者:
Kitney R
Kitney R
中科院分区:
工程技术2区
文献类型:
--
作者:
Casas A;Bultelle M;Motraghi C;Kitney R

文献摘要

被引文献

相似文献

我们提出了一个软件工具,称为cMatch,重建和识别合成的基因结构,从他们的序列,或一组子序列-基于两个实际的信息:他们的模块化结构,和组件库。虽然开发的组合途径工程问题,并解决其质量控制(QC)的瓶颈,cMatch并不限于这些应用。QC在组装、转化和生长后进行。它有一个简单的目标,即验证细胞中包含的遗传物质是否与预期构建的物质相匹配,如果不是,则定位差异并估计其严重性。在再现性/可靠性方面,QC步骤至关重要。在该步骤失败需要重复构建和/或测序步骤。当手动或半手动执行时,QC是一个非常耗时、容易出错的过程,其与构建体的数量及其复杂性的比例非常差。为了使质量控制更加顺畅和可靠,cMatch执行了一个我们称之为“构造匹配”的操作,并将其自动化。构造匹配比简单的序列匹配更彻底,因为它在功能级别进行匹配,并在单个组件级别和整个构造上量化匹配。提出了两种算法(CM_1和CM_2)。它们根据输入的性质而不同。CM_1是用于构建体匹配的核心算法,并且当输入序列足够长以覆盖其整体的构建体(例如,用诸如下一代测序的方法获得)。CM_2是一个扩展,旨在处理较短的数据(例如,通过桑格测序获得),并且需要重组。这两种算法都能在几分钟内(即使在处理能力有限的硬件上)产生准确的构造匹配,以及一组可用于提高决策过程鲁棒性的指标。为了确保可靠性和再现性,cMatch建立在高度验证的成对匹配Smith-Waterman算法之上。所有提出的测试都是在具有挑战性但现实的结构的合成数据上进行的,以及在代谢工程实例(番茄红素生产)研究期间收集的真实的数据上进行的。
We present a software tool, called cMatch, to reconstruct and identify synthetic genetic constructs from their sequences, or a set of sub-sequences—based on two practical pieces of information: their modular structure, and libraries of components. Although developed for combinatorial pathway engineering problems and addressing their quality control (QC) bottleneck, cMatch is not restricted to these applications. QC takes place post assembly, transformation and growth. It has a simple goal, to verify that the genetic material contained in a cell matches what was intended to be built - and when it is not the case, to locate the discrepancies and estimate their severity. In terms of reproducibility/reliability, the QC step is crucial. Failure at this step requires repetition of the construction and/or sequencing steps. When performed manually or semi-manually QC is an extremely time-consuming, error prone process, which scales very poorly with the number of constructs and their complexity. To make QC frictionless and more reliable, cMatch performs an operation we have called “construct-matching” and automates it. Construct-matching is more thorough than simple sequence-matching, as it matches at the functional level-and quantifies the matching at the individual component level and across the whole construct. Two algorithms (called CM_1 and CM_2) are presented. They differ according to the nature of their inputs. CM_1 is the core algorithm for construct-matching and is to be used when input sequences are long enough to cover constructs in their entirety (e.g., obtained with methods such as next generation sequencing). CM_2 is an extension designed to deal with shorter data (e.g., obtained with Sanger sequencing), and that need recombining. Both algorithms are shown to yield accurate construct-matching in a few minutes (even on hardware with limited processing power), together with a set of metrics that can be used to improve the robustness of the decision-making process. To ensure reliability and reproducibility, cMatch builds on the highly validated pairwise-matching Smith-Waterman algorithm. All the tests presented have been conducted on synthetic data for challenging, yet realistic constructs - and on real data gathered during studies on a metabolic engineering example (lycopene production).