PPIT: an R package for inferring microbial taxonomy from nifH sequences

PPIT: an R package for inferring microbial taxonomy from nifH sequences
复制标题

PPIT:用于从 nifH 序列推断微生物分类的 R 包

DOI:
10.1093/bioinformatics/btab100
复制
发表时间:
2021
期刊:
影响因子:
5.8
通讯作者:
Dekas, Anne E
Dekas, Anne E
中科院分区:
生物学3区
文献类型:
--
作者:
Kapili, Bennett J;Dekas, Anne E

文献摘要

被引文献

相似文献

动力将微生物群落成员与其生态功能联系起来是环境微生物学的中心目标。当被指定为分类学时,代谢标记基因的扩增序列可以暗示这种联系,从而提供了支撑特定生态系统功能的系统发育结构的概述。然而,从代谢标记基因序列推断微生物分类仍然是一个挑战,特别是对于经常被测序的固氮标记基因--固氮酶还原酶(NifH)。最近的水平基因转移他的进化历史可能会混淆从现有软件中使用的成对身份识别方法得出的分类推断。其他用于推断分类的方法没有标准化,需要人工检查,难以规模化。结果我们提出了用于推断分类的系统发育放置(PPIT),这是一个R包,使用系统发育和序列同源性方法从nifHanticons推断微生物分类。当用户将查询序列放置在PPIT提供的参考序列树(n= 6317全长H序列)上后,PPIT搜索每个查询序列的系统发育邻域,并尝试推断微生物分类学。只有当系统发育邻域中的参考符合:(1)分类上一致,并且(2)与查询共享足够的成对同一性,从而避免了由于已知的水平基因转移事件而导致的错误推断时,才会得出推断。我们发现,与基于BLAST的方法相比,PPIT以更少的总推理为代价返回了更高比例的正确分类推理。我们在深海沉积物上展示了PPIT,发现DeltaProtebacia是最丰富的潜在重氮菌。使用该数据集,我们证明了基于查询序列放置视觉检查的PPIT推理修正可以实现对查询集中几乎所有序列的分类推理。我们还讨论了用户如何将PPIT应用于其他标记基因的分析。可用性和实现PPIT在https://github.com/bkapili/ppit.上向非商业性用户免费提供安装包括一个演示包使用的小插曲,并重现了这里讨论的分析。RawnifHantons序列数据已保存在GenBank、EMBL和DDBJ数据库中,BioProject编号为PRJEB37167。补充信息补充数据可在BioInformation Online上获得。
MotivationLinking microbial community members to their ecological functions is a central goal of environmental microbiology. When assigned taxonomy, amplicon sequences of metabolic marker genes can suggest such links, thereby offering an overview of the phylogenetic structure underpinning particular ecosystem functions. However, inferring microbial taxonomy from metabolic marker gene sequences remains a challenge, particularly for the frequently sequenced nitrogen fixation marker gene, nitrogenase reductase (nifH). Horizontal gene transfer in recentnifHevolutionary history can confound taxonomic inferences drawn from the pairwise identity methods used in existing software. Other methods for inferring taxonomy are not standardized and require manual inspection that is difficult to scale.ResultsWe present Phylogenetic Placement for Inferring Taxonomy (PPIT), an R package that infers microbial taxonomy fromnifHamplicons using both phylogenetic and sequence identity approaches. After users place query sequences on a referencenifHgene tree provided by PPIT (n= 6317 full-lengthnifHsequences), PPIT searches the phylogenetic neighborhood of each query sequence and attempts to infer microbial taxonomy. An inference is drawn only if references in the phylogenetic neighborhood are: (1) taxonomically consistent and (2) share sufficient pairwise identity with the query, thereby avoiding erroneous inferences due to known horizontal gene transfer events. We find that PPIT returns a higher proportion of correct taxonomic inferences than BLAST-based approaches at the cost of fewer total inferences. We demonstrate PPIT on deep-sea sediment and find thatDeltaproteobacteriaare the most abundant potential diazotrophs. Using this dataset, we show that emending PPIT inferences based on visual inspection of query sequence placement can achieve taxonomic inferences for nearly all sequences in a query set. We additionally discuss how users can apply PPIT to the analysis of other marker genes.Availability and implementationPPIT is freely available to noncommercial users at https://github.com/bkapili/ppit. Installation includes a vignette that demonstrates package use and reproduces thenifHamplicon analysis discussed here. The rawnifHamplicon sequence data have been deposited in the GenBank, EMBL and DDBJ databases under BioProject number PRJEB37167.Supplementary informationSupplementary data are available atBioinformaticsonline.