CALDERA: finding all significant de Bruijn subgraphs for bacterial GWAS.

CALDERA: finding all significant de Bruijn subgraphs for bacterial GWAS.
复制标题

DOI:
10.1093/bioinformatics/btac238
复制
发表时间:
2022-06-24
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

全基因组关联研究旨在发现与某一性状相关的遗传变异,已被广泛应用于细菌,以确定耐药或超强毒力的基因决定因素。最近的细菌GWA法通常依赖于k-MERS,它在基因组中的存在可以表示从单核苷酸多态到可移动遗传元件的各种变异。这种方法不需要参考基因组,因此更容易解释辅助基因。然而,同一基因可能存在于不同菌株之间的略有不同的版本中,导致稀释效应。在这里,我们通过测试由基因组k-mers上定义的de Bruijn图的闭连通子图(CCSS)构建的协变量来克服这个问题。这些协变量将多态基因作为一个单一实体捕获,在能力和可解释性方面都改善了基于k-mer的GWA。然而,单纯地测试所有可能的子图的方法将由于多次测试修正而变得无能为力,而仅仅探索这些子图将很快变得难以计算。可检验假说的概念已经被成功地用于在相似的背景下解决这两个问题。我们利用这一概念对所有CCSS进行测试,为这些对象提出了一种新的枚举方案,该方案充分利用了可测试性提供的剪枝机会,从而大大提高了计算效率。我们的方法与现有的可视化工具相结合,以便于解释。我们提供了我们的方法的实现,以及在https://github.com/HectorRDB/Caldera_ISMB.上重现所有结果的代码补充数据可在生物信息学在线上获得。
Genome-wide association studies (GWAS), aiming to find genetic variants associated with a trait, have widely been used on bacteria to identify genetic determinants of drug resistance or hypervirulence. Recent bacterial GWAS methods usually rely on k-mers, whose presence in a genome can denote variants ranging from single-nucleotide polymorphisms to mobile genetic elements. This approach does not require a reference genome, making it easier to account for accessory genes. However, a same gene can exist in slightly different versions across different strains, leading to diluted effects. Here, we overcome this issue by testing covariates built from closed connected subgraphs (CCSs) of the de Bruijn graph defined over genomic k-mers. These covariates capture polymorphic genes as a single entity, improving k-mer-based GWAS both in terms of power and interpretability. However, a method naively testing all possible subgraphs would be powerless due to multiple testing corrections, and the mere exploration of these subgraphs would quickly become computationally intractable. The concept of testable hypothesis has successfully been used to address both problems in similar contexts. We leverage this concept to test all CCSs by proposing a novel enumeration scheme for these objects which fully exploits the pruning opportunity offered by testability, resulting in drastic improvements in computational efficiency. Our method integrates with existing visual tools to facilitate interpretation. We provide an implementation of our method, as well as code to reproduce all results at https://github.com/HectorRDB/Caldera_ISMB. Supplementary data are available at Bioinformatics online.
DOI: 10.1186/s12864-016-2889-6
发表时间: 2016-09-26
期刊: BMC genomics
影响因子: 4.4
作者:
Drouin A;Giguère S;Déraspe M;Marchand M;Tyers M;Loo VG;Bourgault AM;Laviolette F;Corbeil J
通讯作者: Corbeil J
DOI: 10.1186/s13059-021-02427-7
发表时间: 2021-07-14
期刊: Genome biology
影响因子: 12.3
作者:
Karcher N;Nigro E;Punčochář M;Blanco-Míguez A;Ciciani M;Manghi P;Zolfo M;Cumbo F;Manara S;Golzato D;Cereseto A;Arumugam M;Bui TPN;Tytgat HLP;Valles-Colomer M;de Vos WM;Segata N
通讯作者: Segata N
DOI: 10.2307/2340521
发表时间: 1922-01-01
影响因子: --
作者:
Fisher, RA
通讯作者: Fisher, RA
DOI: 10.1093/bioinformatics/btv263
发表时间: 2015-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Llinares-López F;Grimm DG;Bodenham DA;Gieraths U;Sugiyama M;Rowan B;Borgwardt K
通讯作者: Borgwardt K
DOI: 10.1093/bioinformatics/btx071
发表时间: 2017-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Llinares-López F;Papaxanthos L;Bodenham D;Roqueiro D;COPDGene Investigators;Borgwardt K
通讯作者: Borgwardt K