Computational graph pangenomics: a tutorial on data structures and their applications.

Computational graph pangenomics: a tutorial on data structures and their applications.
复制标题

DOI:
10.1007/s11047-022-09882-6
复制
发表时间:
2022-03
期刊:
影响因子:
2.1
通讯作者:
Siren, Jouni
Siren, Jouni
中科院分区:
计算机科学4区
文献类型:
--
作者:
Baaijens, Jasmijn A.;Bonizzoni, Paola;Boucher, Christina;Della Vedova, Gianluca;Pirola, Yuri;Rizzi, Raffaella;Siren, Jouni

文献摘要

参考文献

被引文献

相似文献

计算泛基因组学是一个新兴的研究领域,它正在改变计算机科学家在生物序列分析中面临的挑战。在过去的几十年里,组合学、弦论、图论和数据结构的贡献对开发大量用于分析人类基因组的软件工具至关重要。这些工具使计算生物学家得以在种群规模上进行雄心勃勃的项目,如1000基因组计划。1000基因组项目的一个主要贡献是描述了人类基因组中广泛的遗传变异,包括在南亚、非洲和欧洲群体中发现了新的变异--从而增强了参考基因组内的变异性目录。目前,在个人化的医学方法中考虑到群体基因组的高度变异性以及个体基因组的特殊性的需要正在迅速推动使用单一参考基因组的传统范式的放弃。一种基于图形的多基因组表示,或称图形基因组,正在取代线性参考基因组。这意味着完全重新考虑分析、存储和获取来自基因组表示的信息的成熟程序。妥善应对这些挑战,对于面对雄心勃勃的医疗项目的计算任务至关重要,这些项目旨在通过对100万人进行测序来表征人类多样性。本教程旨在向读者介绍表示图的数据结构理论的最新进展。我们讨论了单倍型的有效表示和图形盘状体中基因类型的可变性,并强调了在解决人类和微生物(病毒)盘状体中的计算问题方面的应用。
Computational pangenomics is an emerging research field that is changing the way computer scientists are facing challenges in biological sequence analysis. In past decades, contributions from combinatorics, stringology, graph theory and data structures were essential in the development of a plethora of software tools for the analysis of the human genome. These tools allowed computational biologists to approach ambitious projects at population scale, such as the 1000 Genomes Project. A major contribution of the 1000 Genomes Project is the characterization of a broad spectrum of genetic variations in the human genome, including the discovery of novel variations in the South Asian, African and European populations—thus enhancing the catalogue of variability within the reference genome. Currently, the need to take into account the high variability in population genomes as well as the specificity of an individual genome in a personalized approach to medicine is rapidly pushing the abandonment of the traditional paradigm of using a single reference genome. A graph-based representation of multiple genomes, or a graph pangenome, is replacing the linear reference genome. This means completely rethinking well-established procedures to analyze, store, and access information from genome representations. Properly addressing these challenges is crucial to face the computational tasks of ambitious healthcare projects aiming to characterize human diversity by sequencing 1M individuals. This tutorial aims to introduce readers to the most recent advances in the theory of data structures for the representation of graph pangenomes. We discuss efficient representations of haplotypes and the variability of genotypes in graph pangenomes, and highlight applications in solving computational problems in human and microbial (viral) pangenomes.
DOI: 10.1016/j.isci.2019.07.011
发表时间: 2019-08-30
期刊: ISCIENCE
影响因子: 5.8
作者:
Denti, Luca;Previtali, Marco;Bonizzoni, Paola
通讯作者: Bonizzoni, Paola
DOI: 10.1186/s12859-015-0533-0
发表时间: 2015-05-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Vyverman M;Baets BD;Fack V;Dawyndt P
通讯作者: Dawyndt P
DOI: 10.1146/annurev-genom-120219-080406
发表时间: 2020-08-31
影响因子: 8.7
作者:
Eizenga JM;Novak AM;Sibbesen JA;Heumos S;Ghaffaari A;Hickey G;Chang X;Seaman JD;Rounthwaite R;Ebler J;Rautiainen M;Garg S;Paten B;Marschall T;Sirén J;Garrison E
通讯作者: Garrison E
DOI: 10.1186/s13059-019-1774-4
发表时间: 2019-08-09
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Ballouz, Sara;Dobin, Alexander;Gillis, Jesse A.
通讯作者: Gillis, Jesse A.
DOI: 10.1093/bioinformatics/btw279
发表时间: 2016-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Chikhi R;Limasset A;Medvedev P
通讯作者: Medvedev P