Annotation of the Giardia proteome through structure-based homology and machine learning

Annotation of the Giardia proteome through structure-based homology and machine learning
复制标题

DOI:
10.1093/gigascience/giy150
复制
发表时间:
2019-01-01
期刊:
影响因子:
9.2
通讯作者:
Jex, Aaron R.
Jex, Aaron R.
中科院分区:
生物学2区
文献类型:
--
作者:
Ansell, Brendan R. E.;Pope, Bernard J.;Jex, Aaron R.

文献摘要

被引文献

相似文献

背景:蛋白质结构的大规模计算预测是一种替代经验结构确定的经济有效的替代方法,对于非模型生物体和被忽视的病原体具有特别的前景。传统的基于序列的工具不足以注释这种不同的生物系统的基因组。相反,蛋白质结构容忍初级氨基酸序列的巨大变化,因此是生化功能的一个强有力的指示器。结构蛋白质组学有望成为病原体基因组学研究的标准部分;然而,现在需要信息学方法来对大量预测的结构进行置信度分配。目的:我们的目标是预测一种被忽视的人类病原体十二指肠贾第鞭毛虫的蛋白质组,并使用分离和组合的各种度量将预测的结构分层为高置信度和低置信度类别。方法:我们使用I-Tasser套件预测了类似于G5,000个在十二指肠编码的蛋白质的结构模型,并在蛋白质数据库中确定了它们最接近的经验性确定的结构同源物。根据查询和参考肽中匹配的蛋白质家族(Pfam)结构域的存在,模型被分配到高置信度或低置信度类别。通过开发随机森林分类器,评估了从套件输出的指标和派生指标单独预测高置信度类别以及组合预测的能力。结果:我们确定了1095个高置信度模型,包括212个假设蛋白质。查询肽和参考肽之间的氨基酸同一性是高置信度状态的最大个体预测因子;然而,随机森林分类器在隔离方面的表现优于任何度量(接收器操作特征曲线下的面积=0.976),并识别了与假阳性预测相对应的305个高置信度类模型的子集。高置信度模型表现出更大的转录丰度,分类器在不同物种之间通用,表明这种方法在自动分层预测结构方面具有广泛的实用价值。附加的基于结构的聚类被用来交叉检查扩展的Nek激酶家族中的置信度预测。几个高置信度的类蛋白质为十二指肠胃的氧化还原平衡机制提供了实质性的新见解,十二指肠胃肠道氧化还原平衡是有限的抗贾第鞭毛虫药物疗效的核心系统。结论:结构蛋白质组学和机器学习相结合可以帮助对包括人类病原体在内的遗传差异的生物进行基因组注释,并对预测的结构进行分层,以促进有限资源的有效分配用于实验研究。
Background: Large-scale computational prediction of protein structures represents a cost-effective alternative to empirical structure determination with particular promise for non-model organisms and neglected pathogens. Conventional sequence-based tools are insufficient to annotate the genomes of such divergent biological systems. Conversely, protein structure tolerates substantial variation in primary amino acid sequence and is thus a robust indicator of biochemical function. Structural proteomics is poised to become a standard part of pathogen genomics research; however, informatic methods are now required to assign confidence in large volumes of predicted structures. Aims: Our aim was to predict the proteome of a neglected human pathogen, Giardia duodenalis, and stratify predicted structures into high- and lower-confidence categories using a variety of metrics in isolation and combination. Methods: We used the I-TASSER suite to predict structural models for similar to 5,000 proteins encoded in G. duodenalis and identify their closest empirically-determined structural homologues in the Protein Data Bank. Models were assigned to high- or lower-confidence categories depending on the presence of matching protein family (Pfam) domains in query and reference peptides. Metrics output from the suite and derived metrics were assessed for their ability to predict the high-confidence category individually, and in combination through development of a random forest classifier. Results: We identified 1,095 high-confidence models including 212 hypothetical proteins. Amino acid identity between query and reference peptides was the greatest individual predictor of high-confidence status; however, the random forest classifier outperformed any metric in isolation (area under the receiver operating characteristic curve = 0.976) and identified a subset of 305 high-confidence-like models, corresponding to false-positive predictions. High-confidence models exhibited greater transcriptional abundance, and the classifier generalized across species, indicating the broad utility of this approach for automatically stratifying predicted structures. Additional structure-based clustering was used to cross-check confidence predictions in an expanded family of Nek kinases. Several high-confidence-like proteins yielded substantial new insight into mechanisms of redox balance in G. duodenalis-a system central to the efficacy of limited anti-giardial drugs. Conclusion: Structural proteomics combined with machine learning can aid genome annotation for genetically divergent organisms, including human pathogens, and stratify predicted structures to promote efficient allocation of limited resources for experimental investigation.