evSeq: Cost-Effective Amplicon Sequencing of Every Variant in a Protein Library

evSeq: Cost-Effective Amplicon Sequencing of Every Variant in a Protein Library
复制标题

DOI:
10.1021/acssynbio.1c00592
复制
发表时间:
2022-03-18
影响因子:
4.7
通讯作者:
Arnold, Frances H.
Arnold, Frances H.
中科院分区:
生物学2区
文献类型:
--
作者:
Wittmann, Bruce J.;Johnston, Kadina E.;Arnold, Frances H.

文献摘要

被引文献

相似文献

蛋白质序列适应性数据的广泛可用性将彻底改变我们对蛋白质的生化理解以及我们对蛋白质的工程能力。不幸的是,即使在典型的蛋白质工程活动中产生了数千种蛋白质变体并对其进行了适应性评估,但大多数从未被测序,留下了大量潜在的序列适应性信息。首先,这是因为测序对于许多蛋白质工程策略是不必要的;因此测序的额外成本和努力是不合理的。这也是因为,尽管已经开发了许多低成本测序策略,但它们通常需要至少一些测序或计算资源的访问和经验,这两者都可能是访问的障碍。在这里,我们介绍了每个变体测序(evSeq),一种方法和工具/标准化组件的集合,用于对蛋白质工程活动期间产生的每个变体基因内的可变区进行测序,每个变体的成本为美分。evSeq旨在为蛋白质工程师以及任何对工程生物系统感兴趣的人提供低成本测序。其湿实验室组件的执行很简单,不需要测序经验来执行,只依赖于生物实验室通常可用的资源和服务,并巧妙地插入现有的蛋白质工程工作流程。evSeq数据的分析同样通过其附带的软件变得简单(可以在github.com/fhalab/evSeq上找到,文档在fhalab.github.io/evSeq上),该软件可以在个人笔记本电脑上运行,并且旨在让没有计算经验的用户访问。evSeq低成本且易于使用,使得收集广泛的蛋白质变体序列适应性数据变得实用。
Widespread availability of protein sequence-fitness data would revolutionize both our biochemical understanding of proteins and our ability to engineer them. Unfortunately, even though thousands of protein variants are generated and evaluated for fitness during a typical protein engineering campaign, most are never sequenced, leaving a wealth of potential sequence-fitness information untapped. Primarily, this is because sequencing is unnecessary for many protein engineering strategies; the added cost and effort of sequencing are thus unjustified. It also results from the fact that, even though many lower-cost sequencing strategies have been developed, they often require at least some access to and experience with sequencing or computational resources, both of which can be barriers to access. Here, we present every variant sequencing (evSeq), a method and collection of tools/standardized components for sequencing a variable region within every variant gene produced during a protein engineering campaign at a cost of cents per variant. evSeq was designed to democratize low-cost sequencing for protein engineers and, indeed, anyone interested in engineering biological systems. Execution of its wet-lab component is simple, requires no sequencing experience to perform, relies only on resources and services typically available to biology labs, and slots neatly into existing protein engineering workflows. Analysis of evSeq data is likewise made simple by its accompanying software (found at github.com/fhalab/evSeq, documentation at fhalab.github.io/evSeq), which can be run on a personal laptop and was designed to be accessible to users with no computational experience. Low-cost and easy-to-use, evSeq makes the collection of extensive protein variant sequence-fitness data practical.