PPIT: an R package for inferring microbial taxonomy from nifH sequences
PPIT: an R package for inferring microbial taxonomy from nifH sequences
复制标题
PPIT:用于从 nifH 序列推断微生物分类的 R 包
DOI:
10.1093/bioinformatics/btab100
复制
发表时间:
2021
期刊:
影响因子:
5.8
通讯作者:
Dekas, Anne E
中科院分区:
文献类型:
--
作者:
Kapili, Bennett J;Dekas, Anne E
MotivationLinking microbial community members to their ecological functions is a central goal of environmental microbiology. When assigned taxonomy, amplicon sequences of metabolic marker genes can suggest such links, thereby offering an overview of the phylogenetic structure underpinning particular ecosystem functions. However, inferring microbial taxonomy from metabolic marker gene sequences remains a challenge, particularly for the frequently sequenced nitrogen fixation marker gene, nitrogenase reductase (nifH). Horizontal gene transfer in recentnifHevolutionary history can confound taxonomic inferences drawn from the pairwise identity methods used in existing software. Other methods for inferring taxonomy are not standardized and require manual inspection that is difficult to scale.ResultsWe present Phylogenetic Placement for Inferring Taxonomy (PPIT), an R package that infers microbial taxonomy fromnifHamplicons using both phylogenetic and sequence identity approaches. After users place query sequences on a referencenifHgene tree provided by PPIT (n= 6317 full-lengthnifHsequences), PPIT searches the phylogenetic neighborhood of each query sequence and attempts to infer microbial taxonomy. An inference is drawn only if references in the phylogenetic neighborhood are: (1) taxonomically consistent and (2) share sufficient pairwise identity with the query, thereby avoiding erroneous inferences due to known horizontal gene transfer events. We find that PPIT returns a higher proportion of correct taxonomic inferences than BLAST-based approaches at the cost of fewer total inferences. We demonstrate PPIT on deep-sea sediment and find thatDeltaproteobacteriaare the most abundant potential diazotrophs. Using this dataset, we show that emending PPIT inferences based on visual inspection of query sequence placement can achieve taxonomic inferences for nearly all sequences in a query set. We additionally discuss how users can apply PPIT to the analysis of other marker genes.Availability and implementationPPIT is freely available to noncommercial users at https://github.com/bkapili/ppit. Installation includes a vignette that demonstrates package use and reproduces thenifHamplicon analysis discussed here. The rawnifHamplicon sequence data have been deposited in the GenBank, EMBL and DDBJ databases under BioProject number PRJEB37167.Supplementary informationSupplementary data are available atBioinformaticsonline.