SACCHARIS: an automated pipeline to streamline discovery of carbohydrate active enzyme activities within polyspecific families and de novo sequence datasets.

SACCHARIS: an automated pipeline to streamline discovery of carbohydrate active enzyme activities within polyspecific families and de novo sequence datasets.
复制标题

DOI:
10.1186/s13068-018-1027-x
复制
发表时间:
2018
影响因子:
6.3
通讯作者:
Abbott DW
Abbott DW
中科院分区:
工程技术1区
文献类型:
--
作者:
Jones DR;Thomas D;Alger N;Ghavidel A;Inglis GD;Abbott DW

文献摘要

参考文献

被引文献

相似文献

在网上数据库中储存新的基因序列正在以前所未有的速度扩大。因此,序列识别继续超过碳水化合物活性酶(CAZymes)的功能表征。在这种模式下,具有新功能的酶的发现经常受到大量未表征序列的阻碍,特别是当酶序列属于表现出不同功能特异性(即多特异性)的家族时。因此,为了指导基于序列的新酶活性的发现和表征,我们开发了一个自动化的电子流水线:用于快速信息特异性预测的碳水化合物活性酶的序列分析和聚类(SACCHARIS)。这条管道简化了从CAZy网站或用户定义的数据集中目前维护的家族中发现新的CAZyme或CBM特异性的未表征序列的选择。SACCHARIS被用来生成GH43的系统发育树,GH43是一个定义了亚科名称的CAZyme家族。这一分析证实,大型数据集可以被组织成具有相关功能的可管理大小的序列簇。用杜氏杆菌DSM 17855(BdGH43b)中的GH43序列播种这棵树,发现它在树中被分割为单一序列。这一模式与它对GH43具有独特的酶活性一致,因为BdGH43b是首次描述的该家族的α-葡聚糖酶。使用家族6 CBM(即,CBM6s)证明了SACCHARIS提取和聚类特征碳水化合物结合模块序列的能力。该CBM家族显示多特异性配体结合谱,并包含许多结构确定的成员。利用SACCHARIS鉴定了一组不同的序列,从一个独特的分支中发现了一个CBM6序列与酵母甘露聚糖结合,这是首次描述了α-甘露聚糖结合的CBM。此外,我们还对内部测序的细菌基因组进行了CAZome分析,并对B.thaiotaomicron VPI-5482和B.thaiotaomicron 7330进行了比较分析,以证明SACCHARIS可以产生“CAZome指纹”,区分两个相关菌株在硅胶中的糖解潜力。在多个特定的CAZyme家族中建立序列-功能和序列-结构关系是简化酶发现的有前途的方法。SACCHARIS通过将从生化到结构特征序列生成的CAZyme和CBM家系树与具有未知功能的蛋白质序列嵌入,从而促进了这一过程。此外,这些树可以与用户定义的数据集(例如基因组学、元基因组学和转录组学)相集成,以提供当前未进行整理的新CAZyme或CBM的实验特征,并供研究人员比较整个CAZome之间的差异序列模式。在这一点上,SACCHARIS提供了一种电子工具,可以为日益复杂的数据集中的酶生物勘探和糖生物技术中的各种应用量身定做。
Deposition of new genetic sequences in online databases is expanding at an unprecedented rate. As a result, sequence identification continues to outpace functional characterization of carbohydrate active enzymes (CAZymes). In this paradigm, the discovery of enzymes with novel functions is often hindered by high volumes of uncharacterized sequences particularly when the enzyme sequence belongs to a family that exhibits diverse functional specificities (i.e., polyspecificity). Therefore, to direct sequence-based discovery and characterization of new enzyme activities we have developed an automated in silico pipeline entitled: Sequence Analysis and Clustering of CarboHydrate Active enzymes for Rapid Informed prediction of Specificity (SACCHARIS). This pipeline streamlines the selection of uncharacterized sequences for discovery of new CAZyme or CBM specificity from families currently maintained on the CAZy website or within user-defined datasets. SACCHARIS was used to generate a phylogenetic tree of a GH43, a CAZyme family with defined subfamily designations. This analysis confirmed that large datasets can be organized into sequence clusters of manageable sizes that possess related functions. Seeding this tree with a GH43 sequence from Bacteroides dorei DSM 17855 (BdGH43b, revealed it partitioned as a single sequence within the tree. This pattern was consistent with it possessing a unique enzyme activity for GH43 as BdGH43b is the first described α-glucanase described for this family. The capacity of SACCHARIS to extract and cluster characterized carbohydrate binding module sequences was demonstrated using family 6 CBMs (i.e., CBM6s). This CBM family displays a polyspecific ligand binding profile and contains many structurally determined members. Using SACCHARIS to identify a cluster of divergent sequences, a CBM6 sequence from a unique clade was demonstrated to bind yeast mannan, which represents the first description of an α-mannan binding CBM. Additionally, we have performed a CAZome analysis of an in-house sequenced bacterial genome and a comparative analysis of B. thetaiotaomicron VPI-5482 and B. thetaiotaomicron 7330, to demonstrate that SACCHARIS can generate “CAZome fingerprints”, which differentiate between the saccharolytic potential of two related strains in silico. Establishing sequence-function and sequence-structure relationships in polyspecific CAZyme families are promising approaches for streamlining enzyme discovery. SACCHARIS facilitates this process by embedding CAZyme and CBM family trees generated from biochemically to structurally characterized sequences, with protein sequences that have unknown functions. In addition, these trees can be integrated with user-defined datasets (e.g., genomics, metagenomics, and transcriptomics) to inform experimental characterization of new CAZymes or CBMs not currently curated, and for researchers to compare differential sequence patterns between entire CAZomes. In this light, SACCHARIS provides an in silico tool that can be tailored for enzyme bioprospecting in datasets of increasing complexity and for diverse applications in glycobiotechnology.
DOI: 10.1371/journal.pone.0038134
发表时间: 2012
期刊: PloS one
影响因子: 3.7
作者:
Ferrer M;Ghazi A;Beloqui A;Vieites JM;López-Cortés N;Marín-Navarro J;Nechitaylo TY;Guazzaroni ME;Polaina J;Waliczek A;Chernikova TN;Reva ON;Golyshina OV;Golyshin PN
通讯作者: Golyshin PN
DOI: 10.1021/bi1006139
发表时间: 2010-07-27
期刊: BIOCHEMISTRY
影响因子: 2.9
作者:
Correia, Marcia A. S.;Abbott, D. Wade;Gilbert, Harry J.
通讯作者: Gilbert, Harry J.
DOI: 10.1186/1471-2148-12-186
发表时间: 2012-09-20
影响因子: 3.4
作者:
Aspeborg H;Coutinho PM;Wang Y;Brumer H 3rd;Henrissat B
通讯作者: Henrissat B
模因套件:用于发现和搜索的工具。
DOI: 10.1093/nar/gkp335
发表时间: 2009-07
影响因子: 14.9
作者:
Bailey TL;Boden M;Buske FA;Frith M;Grant CE;Clementi L;Ren J;Li WW;Noble WS
通讯作者: Noble WS
DOI: 10.1016/s0022-2836(03)00152-9
发表时间: 2003-03-28
影响因子: 5.6
作者:
Boraston, AB;Notenboom, V;Davies, G
通讯作者: Davies, G