Computational pan-genomics: status, promises and challenges

Computational pan-genomics: status, promises and challenges
复制标题

计算泛基因组学:现状、前景和挑战

DOI:
10.1093/bib/bbw089
复制
发表时间:
2018-01-01
影响因子:
9.5
通讯作者:
Schonhuth, Alexander
Schonhuth, Alexander
中科院分区:
生物学2区
文献类型:
--
作者:
Marschall, Tobias;Marz, Manja;Schonhuth, Alexander

文献摘要

被引文献

相似文献

从人类遗传学和肿瘤学到植物育种、微生物学和病毒学,许多学科都面临着分析快速增长的基因组数量的挑战。以智人为例,在未来几年内,测序的基因组数量将接近数十万。简单地扩大已建立的生物信息学管道将不足以充分利用如此丰富的基因组数据集的全部潜力。相反,我们需要新颖的、性质不同的计算方法和范式。我们将见证计算泛基因组学的快速扩展,这是计算生物学研究的一个新的子领域。在本文中,我们概括了现有的定义,并将泛基因组理解为共同分析或用作参考的任何基因组序列的集合。我们研究了构建和使用泛基因组的现有方法,讨论了未来技术和方法的潜在好处,并从上述生物学学科的有利位置回顾了开放的挑战。作为计算范式转变的一个突出例子,我们特别强调了从参考基因组表示为字符串到表示为图形的过渡。我们概述了这个和其他来自不同应用领域的挑战如何转化为常见的计算问题,指出了相关的生物信息学技术,并确定了计算机科学中的开放问题。通过这篇综述,我们旨在提高人们的认识,即计算泛基因组学的联合方法可以帮助解决当前在各个领域面临的许多问题。
Many disciplines, from human genetics and oncology to plant breeding, microbiology and virology, commonly face the challenge of analyzing rapidly increasing numbers of genomes. In case of Homo sapiens, the number of sequenced genomes will approach hundreds of thousands in the next few years. Simply scaling up established bioinformatics pipelines will not be sufficient for leveraging the full potential of such rich genomic data sets. Instead, novel, qualitatively different computational methods and paradigms are needed. We will witness the rapid extension of computational pan-genomics, a new sub-area of research in computational biology. In this article, we generalize existing definitions and understand a pan-genome as any collection of genomic sequences to be analyzed jointly or to be used as a reference. We examine already available approaches to construct and use pan-genomes, discuss the potential benefits of future technologies and methodologies and review open challenges from the vantage point of the above-mentioned biological disciplines. As a prominent example for a computational paradigm shift, we particularly highlight the transition from the representation of reference genomes as strings to representations as graphs. We outline how this and other challenges from different application domains translate into common computational problems, point out relevant bioinformatics techniques and identify open problems in computer science. With this review, we aim to increase awareness that a joint approach to computational pan-genomics can help address many of the problems currently faced in various domains.