From next-generation resequencing reads to a high-quality variant data set

From next-generation resequencing reads to a high-quality variant data set
复制标题

从下一代重测序读段到高质量变异数据集

DOI:
10.1038/hdy.2016.102
复制
发表时间:
2016-10-19
期刊:
影响因子:
3.900
通讯作者:
S P Pfeifer
S P Pfeifer
中科院分区:
生物学2区
文献类型:
--
作者:
S P Pfeifer

文献摘要

被引文献

相似文献

测序技术通过允许以前所未有的分辨率分析基因组变异而彻底改变了生物学。高通量测序是快速和廉价的,使其可用于广泛的研究课题。然而,产生的数据包含微妙但复杂的错误,偏差和不确定性,这给可靠的变异检测带来了一些统计和计算挑战。为了充分挖掘高通量测序的潜力,需要对所产生的数据以及可用的方法学进行透彻的理解。在这里,我回顾了几种常用的生成和处理下一代重测序数据的方法,讨论了误差和偏差的影响以及它们对下游分析的影响,并通过突出几种代表当前技术水平的复杂的基于参考的方法,为从原始读数中生成高质量的单核苷酸多态性数据集提供了一般指导和建议。NGS技术的进步以及成本的大幅降低使得高通量测序能够用于广泛的研究课题,包括对多种生物体中遗传多样性的群体规模研究。从获得的序列生成高质量的变异和基因型调用集是一项重要的任务,由于数据中许多微妙但复杂类型的错误,偏差和不确定性而变得复杂。然而,已经提出了许多统计方法和计算工具来应对这些挑战。在这里,我概述了优化原始重测序数据中变异和基因型识别准确性所需的SNP识别工作流程的主要分析步骤。由于可用的软件快速发展以应对不断变化的测序技术和方案,我指出了在选择适合特定研究设计的工具时需要考虑的重要因素,作为有兴趣生成自己的变异数据集的研究人员的指南,而不是推荐一个单一的已建立的变异调用管道。然而,仍然存在挑战。特别是,在高度多样化的物种或基因组区域,或在重复的低复杂性区域的变异的正确识别,仍然是困难的。在这些情况下,有价值的替代方案是无参考的、基于汇编的变体调用(例如,Cortex(Iqbal等人,2012)),其对变体类型和与参考序列的偏离都是不可知的,代价是通常较低的灵敏度和较高的计算要求。此外,单分子、实时、第三代测序和作图技术可以帮助检测这些情况下的变异。虽然尚未完全建立,但它们已经成功地应用于研究遗传多样性(参见,例如; Chaisson等人,2015 ; Gordon等人,2016年),并确实显示了未来基因组研究的巨大前景。
Sequencing has revolutionized biology by permitting the analysis of genomic variation at an unprecedented resolution. High-throughput sequencing is fast and inexpensive, making it accessible for a wide range of research topics. However, the produced data contain subtle but complex types of errors, biases and uncertainties that impose several statistical and computational challenges to the reliable detection of variants. To tap the full potential of high-throughput sequencing, a thorough understanding of the data produced as well as the available methodologies is required. Here, I review several commonly used methods for generating and processing next-generation resequencing data, discuss the influence of errors and biases together with their resulting implications for downstream analyses and provide general guidelines and recommendations for producing high-quality single-nucleotide polymorphism data sets from raw reads by highlighting several sophisticated reference-based methods representing the current state of the art.Over the past years, advances in NGS technologies together with a considerable decrease in costs has permitted high-throughput sequencing to become accessible for a wide range of research topics, including population-scale studies of genetic diversity in a multitude of organisms. Generating high-quality variant and genotype call sets from the obtained sequences is a nontrivial task, complicated by many subtle but complex types of errors, biases and uncertainties in the data. However, many statistical methods and computational tools have been proposed to tackle these challenges. Here, I outlined the main analytic steps of a SNP calling workflow required to optimize the accuracy of variant and genotype calling from raw resequencing data. Because of the fact that available software rapidly evolves in response to the changing sequencing technologies and protocols, I pointed out important factors to consider when choosing tools suitable for a particular study design as a guideline for researchers interested in generating their own variant data sets, rather than recommending a single established variant calling pipeline. Nevertheless, there are still challenges. In particular, the correct identification of variants in highly diverse species or genomic regions, or in repetitive low-complexity regions, remains difficult. Under these circumstances, a valuable alternative is reference-free, assembly-based variant calling (for example, Cortex ( Iqbal et al., 2012 )) that is agnostic to both variant type and divergence from the reference sequence, at the cost of a generally lower sensitivity and higher computational requirements. In addition, single-molecule, real-time, third-generation sequencing and mapping technologies can aid the detection of variation in these cases. Although not yet as well established, they have already been successfully applied to study genetic diversity (see, for example; Chaisson et al., 2015 ; Gordon et al., 2016 ) and indeed show great promise for future genomic research.