High-Specificity Targeted Functional Profiling in Microbial Communities with ShortBRED.

High-Specificity Targeted Functional Profiling in Microbial Communities with ShortBRED.
复制标题

DOI:
10.1371/journal.pcbi.1004557
复制
发表时间:
2015-12
影响因子:
4.3
通讯作者:
Huttenhower C
Huttenhower C
中科院分区:
生物学2区
文献类型:
--
作者:
Kaminski J;Gibson MK;Franzosa EA;Segata N;Dantas G;Huttenhower C

文献摘要

被引文献

相似文献

从宏基因组测序数据分析微生物群落功能仍然是计算上具有挑战性的问题。将来自此类样本的数百万个DNA读段映射到参考蛋白质数据库需要较长的运行时间,并且短的读段长度可能导致对不相关蛋白质的虚假命中(特异性损失)。我们开发了ShortBred(Short,Better Representative Extract Dataset)来应对这些挑战,促进宏基因组样本的快速,准确的功能分析。ShortBRED由两个部分组成:(i)将感兴趣的参考蛋白减少为短的、高度代表性的氨基酸序列(“标记物”)的方法和(ii)将读段映射到这些标记物以定量其相关蛋白质的相对丰度的搜索步骤。在对合成数据进行ShortBRED评估后,我们将其应用于描述来自美国,中国,马拉维和委内瑞拉的肠道微生物组中的抗生素耐药蛋白家族。我们的研究结果支持抗生素耐药性是人类肠道微生物组的核心功能,四环素耐药核糖体保护蛋白和A类β-内酰胺酶是全球分布最广泛的耐药机制。ShortBRED标记适用于其他基于同源性的搜索任务,我们在这里通过识别3,000多个微生物分离株基因组中抗生素耐药性的系统发育特征来证明这一点。ShortBRED可用于分析各种感兴趣的蛋白质家族;软件、源代码和文档可在http://huttenhower.sph.harvard.edu/shortbred下载。人类微生物组的高通量DNA测序为有兴趣研究微生物群落功能(如抗生素耐药性)的研究人员提供了巨大的资源。然而,将DNA读段分配给蛋白质家族仍然是一个具有挑战性的问题,因为来自给定蛋白质编码基因的读段可能虚假地映射到来自不相关蛋白质的同源区域,这导致假阳性。我们用我们的方法ShortBred解决了这个问题,该方法首先识别出对特定蛋白质家族具有高度代表性的短肽序列(“标记”),然后在宏基因组测序数据中搜索这些标记,以准确检测和定量蛋白质家族。在这项工作中,我们应用ShortBred来分析全球健康人类微生物组和细菌基因组中的抗生素耐药性。ShortBRED可以类似地应用于分析许多其他感兴趣的蛋白质家族。
Profiling microbial community function from metagenomic sequencing data remains a computationally challenging problem. Mapping millions of DNA reads from such samples to reference protein databases requires long run-times, and short read lengths can result in spurious hits to unrelated proteins (loss of specificity). We developed ShortBRED (Short, Better Representative Extract Dataset) to address these challenges, facilitating fast, accurate functional profiling of metagenomic samples. ShortBRED consists of two components: (i) a method that reduces reference proteins of interest to short, highly representative amino acid sequences (“markers”) and (ii) a search step that maps reads to these markers to quantify the relative abundance of their associated proteins. After evaluating ShortBRED on synthetic data, we applied it to profile antibiotic resistance protein families in the gut microbiomes of individuals from the United States, China, Malawi, and Venezuela. Our results support antibiotic resistance as a core function in the human gut microbiome, with tetracycline-resistant ribosomal protection proteins and Class A beta-lactamases being the most widely distributed resistance mechanisms worldwide. ShortBRED markers are applicable to other homology-based search tasks, which we demonstrate here by identifying phylogenetic signatures of antibiotic resistance across more than 3,000 microbial isolate genomes. ShortBRED can be applied to profile a wide variety of protein families of interest; the software, source code, and documentation are available for download at http://huttenhower.sph.harvard.edu/shortbred High throughput DNA sequencing of the human microbiome presents a tremendous resource for researchers interested in studying microbial community functions such as antibiotic resistance. However, assigning DNA reads to protein families remains a challenging problem, as reads derived from a given protein-coding gene may map spuriously to homologous regions from unrelated proteins, which results in false positives. We addressed this problem with our method ShortBRED, which first identifies short peptide sequences (“markers”) that are highly representative for specific protein families, and then searches for these markers in metagenomic sequencing data to accurately detect and quantify protein families. In this work, we applied ShortBRED to profile antibiotic resistance in the healthy human microbiome of individuals worldwide and across bacterial genomes. ShortBRED can be similarly applied to profile many other protein families of interest.