De Novo Clustering of Long-Read Transcriptome Data Using a Greedy, Quality Value-Based Algorithm

De Novo Clustering of Long-Read Transcriptome Data Using a Greedy, Quality Value-Based Algorithm
复制标题

DOI:
10.1089/cmb.2019.0299
复制
发表时间:
2020-03-16
影响因子:
1.7
通讯作者:
Medvedev, Paul
Medvedev, Paul
中科院分区:
生物学4区
文献类型:
--
作者:
Sahlin, Kristoffer;Medvedev, Paul

文献摘要

被引文献

相似文献

使用Pacific Biosciences(PacBio)Iso-Seq和Oxford Nanopore Technologies对转录本进行长读序测序已被证明是许多生物体中复杂同种型景观研究的核心。然而,目前的从头转录重建算法从长读数据是有限的,留下的潜力,这些技术未实现。一个常见的瓶颈是缺乏可扩展的和准确的算法,根据其基因家族的起源聚类长读段。为了应对这一挑战,我们开发了isONcloud,这是一种贪婪的聚类算法(规模),并利用质量值(处理可变的错误率)。我们在三个模拟数据集和五个生物数据集上测试了isONcraft,涵盖了生物体,技术和读取深度。我们的研究结果表明,isONcraft是一个实质性的改进,无论是在整体准确性和/或可扩展性的大型数据集。
Long-read sequencing of transcripts with Pacific Biosciences (PacBio) Iso-Seq and Oxford Nanopore Technologies has proven to be central to the study of complex isoform landscapes in many organisms. However, current de novo transcript reconstruction algorithms from long-read data are limited, leaving the potential of these technologies unfulfilled. A common bottleneck is the dearth of scalable and accurate algorithms for clustering long reads according to their gene family of origin. To address this challenge, we develop isONclust, a clustering algorithm that is greedy (to scale) and makes use of quality values (to handle variable error rates). We test isONclust on three simulated and five biological data sets, across a breadth of organisms, technologies, and read depths. Our results demonstrate that isONclust is a substantial improvement over previous approaches, both in terms of overall accuracy and/or scalability to large data sets.