Pacybara: Accurate long-read sequencing for barcoded mutagenized allelic libraries.

Pacybara: Accurate long-read sequencing for barcoded mutagenized allelic libraries.
复制标题

Pacybara:针对条形码诱变等位基因库的准确长读长测序。

DOI:
10.1101/2023.02.22.529427
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Roth,FrederickP
Roth,FrederickP
中科院分区:
--
文献类型:
--
作者:
Weile,Jochen;Ferra,Gabrielle;Boyle,Gabriel;Pendyala,Sriram;Amorosi,Clara;Yeh,Chiann-Ling;Cote,AtinaG;Kishore,Nishka;Tabet,Daniel;vanLoggerenberg,Warren;Rayhan,Ashyad;Fowler,DouglasM;Dunham,MaitreyaJ;Roth,FrederickP

文献摘要

相似文献

长读段测序技术是许多应用的一种有吸引力的解决方案,但通常存在较高的错误率。多个读段的比对可以提高碱基识别准确性,但是一些应用,例如其中多个不同克隆相差一个或几个变体的测序诱变文库,需要使用条形码或独特的分子标识符。不幸的是,测序错误可以干扰正确的条形码识别,和一个给定的条形码序列可以连接到多个独立的克隆在一个给定library.ResultsHere,我们专注于目标应用程序的测序诱变库的背景下,多路复用分析的变体效应(MAVEs)。MAVE越来越多地用于创建全面的基因型-表型图谱,以帮助临床变异解释。许多MAVE方法使用条形码化突变体文库的长读段测序,用于条形码与基因型的准确关联。现有的长读段测序管道不能解释不准确的测序或非独特的条形码。在这里,我们描述了Pacybara,它通过基于(易错)条形码的相似性聚类长读段来处理这些问题,同时还检测与多种基因型相关的条形码。Pacybara还检测重组(嵌合)克隆并减少假阳性indel调用。在三个示例应用程序中,我们展示了Pacybara识别并正确解决了这些问题。可用性和实现Pacybara,可在https://github.com/rothlab/pacybara免费获得,使用R,Python和bash为Linux实现。它通过Slurm、PBS或GridEngine编译器在GNU/Linux HPC集群上运行。也可提供单机单工版本。
MotivationLong-read sequencing technologies, an attractive solution for many applications, often suffer from higher error rates. Alignment of multiple reads can improve base-calling accuracy, but some applications, e.g. sequencing mutagenized libraries where multiple distinct clones differ by one or few variants, require the use of barcodes or unique molecular identifiers. Unfortunately, sequencing errors can interfere with correct barcode identification, and a given barcode sequence may be linked to multiple independent clones within a given library.ResultsHere we focus on the target application of sequencing mutagenized libraries in the context of multiplexed assays of variant effects (MAVEs). MAVEs are increasingly used to create comprehensive genotype-phenotype maps that can aid clinical variant interpretation. Many MAVE methods use long-read sequencing of barcoded mutant libraries for accurate association of barcode with genotype. Existing long-read sequencing pipelines do not account for inaccurate sequencing or nonunique barcodes. Here, we describe Pacybara, which handles these issues by clustering long reads based on the similarities of (error-prone) barcodes while also detecting barcodes that have been associated with multiple genotypes. Pacybara also detects recombinant (chimeric) clones and reduces false positive indel calls. In three example applications, we show that Pacybara identifies and correctly resolves these issues.Availability and implementationPacybara, freely available at https://github.com/rothlab/pacybara, is implemented using R, Python, and bash for Linux. It runs on GNU/Linux HPC clusters via Slurm, PBS, or GridEngine schedulers. A single-machine simplex version is also available.