The choice of sequence homologs included in multiple sequence alignments has a dramatic impact on evolutionary conservation analysis

The choice of sequence homologs included in multiple sequence alignments has a dramatic impact on evolutionary conservation analysis
复制标题

DOI:
10.1093/bioinformatics/bty523
复制
发表时间:
2019-01-01
期刊:
影响因子:
5.8
通讯作者:
Fiser, Andras
Fiser, Andras
中科院分区:
生物学3区
文献类型:
--
作者:
Gil, Nelson;Fiser, Andras

文献摘要

被引文献

相似文献

动机:半个多世纪以来,序列保守模式的分析已被广泛用于识别功能上重要的(催化和配体结合)蛋白质残基。尽管经过了数十年的发展,平均而言,最先进的非基于模板的功能残基预测方法必须预测大约 25% 的蛋白质总残基,才能正确识别一半的蛋白质功能位点残基。绝大多数误报导致报告的“F 分数”接近 0.3。我们研究了当前方法的局限性,重点关注迄今为止被忽视的多重序列比对 (MSA) 中同源物特定选择的影响。结果:通过调查 1023 个蛋白质的结合位点,探索了基于保守的功能残基预测的局限性。对由从 PSI-BLAST 搜索中随机选择的同系物组成的 MSA 进行简单的保守分析,可实现类似于 0.3 的平均 F 分数,这是由最先进的方法报告的性能匹配,这些方法通常会考虑在机器学习设置中进行预测的附加特征。有趣的是,我们发现简单的组合 MSA 采样算法几乎在每种情况下都会产生具有一组最佳同系物的 MSA,其守恒分析达到类似于 0.6 的平均 F 分数,使最先进的性能提高一倍。我们还表明,考虑到不同结合位点定义之间的一致性,这几乎达到了可能性能的理论极限。此外,我们还展示了最大互信息比对选择 (SAMMI) 在这个方向上取得的进展,这是一种基于信息论的方法,用于识别具有生物信息的 MSA。这项工作强调了优化组成的 MSA 在保护分析中的重要性和未利用的潜力。
Motivation: The analysis of sequence conservation patterns has been widely utilized to identify functionally important (catalytic and ligand-binding) protein residues for over a half-century. Despite decades of development, on average state-of-the-art non-template-based functional residue prediction methods must predict similar to 25% of a protein's total residues to correctly identify half of the protein's functional site residues. The overwhelming proportion of false positives results in reported 'F-Scores' of similar to 0.3. We investigated the limits of current approaches, focusing on the so-far neglected impact of the specific choice of homologs included in multiple sequence alignments (MSAs).Results: The limits of conservation-based functional residue prediction were explored by surveying the binding sites of 1023 proteins. A straightforward conservation analysis of MSAs composed of randomly selected homologs sampled from a PSI-BLAST search achieves average F-Scores of similar to 0.3, a performance matching that reported by state-of-the-art methods, which often consider additional features for the prediction in a machine learning setting. Interestingly, we found that a simple combinatorial MSA sampling algorithm will in almost every case produce an MSA with an optimal set of homologs whose conservation analysis reaches average F-Scores of similar to 0.6, doubling state-of-the-art performance. We also show that this is nearly at the theoretical limit of possible performance given the agreement between different binding site definitions. Additionally, we showcase the progress in this direction made by Selection of Alignment by Maximal Mutual Information (SAMMI), an information-theory-based approach to identifying biologically informative MSAs. This work highlights the importance and the unused potential of optimally composed MSAs for conservation analysis.