Comparative study of the effectiveness and limitations of current methods for detecting sequence coevolution.

Comparative study of the effectiveness and limitations of current methods for detecting sequence coevolution.
复制标题

比较研究当前方法检测序列协同进化的有效性和局限性。

DOI:
10.1093/bioinformatics/btv103
复制
发表时间:
2015-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Bahar I
Bahar I
中科院分区:
其他
文献类型:
--
作者:
Mao W;Kaya C;Dutta A;Horovitz A;Bahar I

文献摘要

参考文献

被引文献

相似文献

研究动机:随着多物种序列数据的快速积累,从多序列比对中提取合理、系统的信息变得越来越重要。目前,有大量的计算方法用于研究沿氨基酸序列成对位置的耦合进化变化,并对结构和功能进行推断。然而,共同进化信号的意义仍有待确定。此外,大量的假阳性(FPs)是由MSA大小不足、系统发育背景和间接耦合引起的。结果:在这里,一组16对非相互作用的蛋白质被彻底检查,以评估不同方法的有效性和局限性。分析表明,最近设计用于消除间接耦合偏差的计算昂贵的方法在检测三级结构接触和消除分子间FPs方面优于其他方法;而互信息等传统方法在提高效率的同时,也得益于洗牌等改进。对来自Negatome数据库的2330对蛋白质家族进行了重复计算,证实了这些结果。最后,使用162个蛋白质家族的训练数据集,我们提出了一种优于现有单个方法的组合方法。总体而言,该研究提供了基于可用MSA大小和计算资源选择合适方法和策略的简单指南。可用性和实现:软件可以通过ProDy API的Evol组件免费获得。补充信息:补充数据可在Bioinformatics在线获取。
Motivation: With rapid accumulation of sequence data on several species, extracting rational and systematic information from multiple sequence alignments (MSAs) is becoming increasingly important. Currently, there is a plethora of computational methods for investigating coupled evolutionary changes in pairs of positions along the amino acid sequence, and making inferences on structure and function. Yet, the significance of coevolution signals remains to be established. Also, a large number of false positives (FPs) arise from insufficient MSA size, phylogenetic background and indirect couplings. Results: Here, a set of 16 pairs of non-interacting proteins is thoroughly examined to assess the effectiveness and limitations of different methods. The analysis shows that recent computationally expensive methods designed to remove biases from indirect couplings outperform others in detecting tertiary structural contacts as well as eliminating intermolecular FPs; whereas traditional methods such as mutual information benefit from refinements such as shuffling, while being highly efficient. Computations repeated with 2,330 pairs of protein families from the Negatome database corroborated these results. Finally, using a training dataset of 162 families of proteins, we propose a combined method that outperforms existing individual methods. Overall, the study provides simple guidelines towards the choice of suitable methods and strategies based on available MSA size and computing resources. Availability and implementation: Software is freely available through the Evol component of ProDy API. Contact: bahar@pitt.edu Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1016/j.cell.2012.04.012
发表时间: 2012-06-22
期刊: Cell
影响因子: 64.5
作者:
Hopf TA;Colwell LJ;Sheridan R;Rost B;Sander C;Marks DS
通讯作者: Marks DS
DOI: 10.1093/bioinformatics/btu458
发表时间: 2014-09-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Michel M;Hayat S;Skwark MJ;Sander C;Marks DS;Elofsson A
通讯作者: Elofsson A
DOI: 10.1093/bioinformatics/btm604
发表时间: 2008-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Dunn, S. D.;Wahl, L. M.;Gloor, G. B.
通讯作者: Gloor, G. B.
DOI: 10.1103/physreve.87.012707
发表时间: 2013-01-11
期刊: PHYSICAL REVIEW E
影响因子: 2.4
作者:
Ekeberg, Magnus;Lovkvist, Cecilia;Aurell, Erik
通讯作者: Aurell, Erik
DOI: 10.1073/pnas.1111471108
发表时间: 2011-12-06
影响因子: 11.1
作者:
Morcos, Faruck;Pagnani, Andrea;Weigt, Martin
通讯作者: Weigt, Martin