Identification of Protein Secretion Systems in Bacterial Genomes Using MacSyFinder

Identification of Protein Secretion Systems in Bacterial Genomes Using MacSyFinder
复制标题

DOI:
10.1007/978-1-4939-7033-9_1
复制
发表时间:
2017-01-01
期刊:
BACTERIAL PROTEIN SECRETION SYSTEMS
影响因子:
--
通讯作者:
Rocha, Eduardo P. C.
Rocha, Eduardo P. C.
中科院分区:
其他
文献类型:
--
作者:
Abby, Sophie S.;Rocha, Eduardo P. C.

文献摘要

被引文献

相似文献

蛋白质分泌系统是复杂的分子机器,其将蛋白质转运通过外膜,并且有时通过多个其他屏障。它们是通过与其他细胞相关的细胞机制的组分的协同选择而进化的,这使得它们有时难以识别和区分。在这里,我们描述了如何识别蛋白质分泌系统的细菌基因组中使用MacSyrup。这种灵活的计算工具使用来自实验研究的知识来识别基因组数据中的同源系统。它可以与一组预定义的模型-“TXSScan”-一起使用,以鉴定双胚层细菌的所有主要分泌系统(即,具有内膜和含有LPS的外膜)。为此,它识别和集群共定位的分泌系统的组成部分,使用序列相似性搜索与隐马尔可夫模型蛋白质谱。最后,它检查集群的遗传内容和组织是否满足模型的约束。TXSScan模型可以自定义以搜索已知系统的变体。这些模型也可以从头开始构建,以识别新的系统。在这一章中,我们描述了一个完整的分析管道,包括一组参考实验研究系统的识别,组件的识别和构建其蛋白质谱,模型的定义,它们的优化,最后,它们作为工具来搜索基因组数据。
Protein secretion systems are complex molecular machineries that translocate proteins through the outer membrane, and sometimes through multiple other barriers. They have evolved by co-option of components from other envelope-associated cellular machineries, making them sometimes difficult to identify and discriminate. Here, we describe how to identify protein secretion systems in bacterial genomes using MacSyFinder. This flexible computational tool uses the knowledge stemming from experimental studies to identify homologous systems in genome data. It can be used with a set of predefined models-"TXSScan"-to identify all major secretion systems of diderm bacteria (i.e., with inner and with LPS-containing outer membranes). For this, it identifies and clusters colocalized components of secretion systems using sequence similarity searches with hidden Markov model protein profiles. Finally, it checks whether the genetic content and organization of clusters satisfy the constraints of the model. TXSScan models can be customized to search for variants of known systems. The models can also be built from scratch to identify novel systems. In this chapter, we describe a complete pipeline of analysis, including the identification of a reference set of experimentally studied systems, the identification of components and the construction of their protein profiles, the definition of the models, their optimization, and, finally, their use as tools to search genomic data.