A novel algorithm for detecting multiple covariance and clustering of biological sequences.

A novel algorithm for detecting multiple covariance and clustering of biological sequences.
复制标题

一种用于检测生物序列的多重协方差和聚类的新算法。

DOI:
10.1038/srep30425
复制
发表时间:
2016-07-25
期刊:
影响因子:
4.6
通讯作者:
Li Y
Li Y
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Shen W;Li Y

文献摘要

被引文献

相似文献

单一的基因突变总是伴随着一组补偿性突变。因此,生物序列通常会发生多重变化,并在维持构象和功能稳定方面起着至关重要的作用。虽然有许多方法可用于检测单个突变或协变对,但检测序列中不同位置的非同步多个变化仍然具有挑战性。在这里,我们开发了一种新的算法,称为Fastcov,它使用独立的配对模型和基于相互限制思想的位点-残基元件串联模型来识别生物序列中的多个相关变化。Fastcov在收获共对和检测多个协变模式方面表现得非常好。通过使用不同尺度的数据集进行10倍交叉验证,特征模式成功地将序列分类为目标组,准确率超过98%。此外,我们证明了多个协变模式代表了与系统发育树相对应的共同进化模式,并提供了对蛋白质结构稳定性的新理解。与其他方法相比,Fastcov不仅提供了一种可靠而有效的方法来识别协变对,而且还具有更强大的功能,包括多重协方差检测和序列分类,对于研究自然选择、药物诱导、环境压力等引起的点突变和补偿性突变最有用。
Single genetic mutations are always followed by a set of compensatory mutations. Thus, multiple changes commonly occur in biological sequences and play crucial roles in maintaining conformational and functional stability. Although many methods are available to detect single mutations or covariant pairs, detecting non-synchronous multiple changes at different sites in sequences remains challenging. Here, we develop a novel algorithm, named Fastcov, to identify multiple correlated changes in biological sequences using an independent pair model followed by a tandem model of site-residue elements based on inter-restriction thinking. Fastcov performed exceptionally well at harvesting co-pairs and detecting multiple covariant patterns. By 10-fold cross-validation using datasets of different scales, the characteristic patterns successfully classified the sequences into target groups with an accuracy of greater than 98%. Moreover, we demonstrated that the multiple covariant patterns represent co-evolutionary modes corresponding to the phylogenetic tree, and provide a new understanding of protein structural stability. In contrast to other methods, Fastcov provides not only a reliable and effective approach to identify covariant pairs but also more powerful functions, including multiple covariance detection and sequence classification, that are most useful for studying the point and compensatory mutations caused by natural selection, drug induction, environmental pressure, etc.