FASTQSim: platform-independent data characterization and in silico read generation for NGS datasets.

FASTQSim: platform-independent data characterization and in silico read generation for NGS datasets.
复制标题

DOI:
10.1186/1756-0500-7-533
复制
发表时间:
2014-08-15
期刊:
影响因子:
1.8
通讯作者:
Shcherbina A
Shcherbina A
中科院分区:
其他
文献类型:
--
作者:
Shcherbina A

文献摘要

被引文献

相似文献

高通量下一代测序技术已经能够快速表征临床和环境样品。因此,可操作数据的最大瓶颈已经成为样本处理和生物信息学分析,需要准确和快速的算法来处理遗传数据。完美的silico数据集的特点是一个有用的工具,用于评估这些算法的性能。在测序的微生物混合物中观察到背景污染微生物。计算机模拟样本提供了准确的事实。为了创建用于评估算法的最佳值,计算机模拟数据应尽可能接近地模拟实际测序仪数据。FASTQSim是一种提供NGS数据集表征和宏基因组数据生成双重功能的工具。FASTQSim是测序平台无关的,并计算任何测序平台的读取长度、质量分数、插入缺失率、单点突变率、插入缺失大小和类似统计数据的分布。为了创建训练或测试数据集,FASTQSim能够将靶序列转换为具有在表征步骤中获得的特定错误特征的计算机模拟读数。FASTQSim使用户能够评估NGS数据集的质量。该工具提供有关读段长度、读段质量、重复和非重复indel概况以及单碱基对置换的信息。FASTQSim允许用户模拟单个读段数据集,这些数据集可用作计划测序项目或基准宏基因组软件的标准化测试场景。在这方面,使用FASTQsim工具生成的计算机数据集与自然数据集相比具有几个优势:它们独立于测序平台,非常好地表征,并且生成成本较低。这些数据集在许多应用中是有价值的,包括多平台汇编程序的培训,基准生物信息学算法性能,以及创建用于检测基因工程工具标记的挑战数据集等。本文的在线版本(doi:10.1186/1756-0500-7-533)包含补充材料,可供授权用户使用。
High-throughput next generation sequencing technologies have enabled rapid characterization of clinical and environmental samples. Consequently, the largest bottleneck to actionable data has become sample processing and bioinformatics analysis, creating a need for accurate and rapid algorithms to process genetic data. Perfectly characterized in silico datasets are a useful tool for evaluating the performance of such algorithms. Background contaminating organisms are observed in sequenced mixtures of organisms. In silico samples provide exact truth. To create the best value for evaluating algorithms, in silico data should mimic actual sequencer data as closely as possible. FASTQSim is a tool that provides the dual functionality of NGS dataset characterization and metagenomic data generation. FASTQSim is sequencing platform-independent, and computes distributions of read length, quality scores, indel rates, single point mutation rates, indel size, and similar statistics for any sequencing platform. To create training or testing datasets, FASTQSim has the ability to convert target sequences into in silico reads with specific error profiles obtained in the characterization step. FASTQSim enables users to assess the quality of NGS datasets. The tool provides information about read length, read quality, repetitive and non-repetitive indel profiles, and single base pair substitutions. FASTQSim allows the user to simulate individual read datasets that can be used as standardized test scenarios for planning sequencing projects or for benchmarking metagenomic software. In this regard, in silico datasets generated with the FASTQsim tool hold several advantages over natural datasets: they are sequencing platform independent, extremely well characterized, and less expensive to generate. Such datasets are valuable in a number of applications, including the training of assemblers for multiple platforms, benchmarking bioinformatics algorithm performance, and creating challenge datasets for detecting genetic engineering toolmarks, etc. The online version of this article (doi:10.1186/1756-0500-7-533) contains supplementary material, which is available to authorized users.