Aligning biological sequences by exploiting residue conservation and coevolution

Aligning biological sequences by exploiting residue conservation and coevolution
复制标题

DOI:
10.1103/physreve.102.062409
复制
发表时间:
2020-12-07
期刊:
影响因子:
2.4
通讯作者:
Zamponi, Francesco
Zamponi, Francesco
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Muntoni, Anna Paola;Pagnani, Andrea;Zamponi, Francesco

文献摘要

被引文献

相似文献

核苷酸序列(对于DNA和RNA)或氨基酸序列(对于蛋白质)是生物学的中心对象。其中最重要的计算问题是序列比对,即,将来自不同生物体的序列以这样的方式排列,以识别相似区域,检测序列之间的进化关系,并预测生物分子的结构和功能。这通常通过概况模型来解决,概况模型捕获位置特异性,如序列中的保守性,但假设不同位置的独立进化。近年来,已经很好地建立了不同氨基酸位置的共同进化对于维持三维结构和功能是必不可少的。基于逆统计物理的建模方法可以捕捉序列集成中的协同进化信号,并且它们现在被广泛用于预测蛋白质结构、蛋白质-蛋白质相互作用和突变景观。在这里,我们提出了DCAlign,一个有效的比对算法的基础上的一个近似的消息传递策略,这是能够克服配置文件模型的局限性,包括位置之间的共同进化,在一般的方式,并因此普遍适用于蛋白质和RNA序列比对,而不需要使用互补的结构信息。使用良好控制的模拟数据以及真实的蛋白质和RNA序列仔细探索DCAlign的潜力。
Sequences of nucleotides (for DNA and RNA) or amino acids (for proteins) are central objects in biology. Among the most important computational problems is that of sequence alignment, i.e., arranging sequences from different organisms in such a way to identify similar regions, to detect evolutionary relationships between sequences, and to predict biomolecular structure and function. This is typically addressed through profile models, which capture position specificities like conservation in sequences but assume an independent evolution of different positions. Over recent years, it has been well established that coevolution of different amino-acid positions is essential for maintaining three-dimensional structure and function. Modeling approaches based on inverse statistical physics can catch the coevolution signal in sequence ensembles, and they are now widely used in predicting protein structure, protein-protein interactions, and mutational landscapes. Here, we present DCAlign, an efficient alignment algorithm based on an approximate message-passing strategy, which is able to overcome the limitations of profile models, to include coevolution among positions in a general way, and to be therefore universally applicable to protein- and RNA-sequence alignment without the need of using complementary structural information. The potential of DCAlign is carefully explored using well-controlled simulated data, as well as real protein and RNA sequences.