Adaptive GDDA-BLAST: fast and efficient algorithm for protein sequence embedding.

Adaptive GDDA-BLAST: fast and efficient algorithm for protein sequence embedding.
复制标题

DOI:
10.1371/journal.pone.0013596
复制
发表时间:
2010-10-22
期刊:
影响因子:
3.7
通讯作者:
van Rossum DB
van Rossum DB
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hong Y;Kang J;Lee D;van Rossum DB

文献摘要

参考文献

被引文献

相似文献

基因组时代的一个主要计算挑战是对现在可用的大量序列信息进行结构/功能注释。大多数蛋白质缺乏全面的注释,这一事实说明了这个问题,即使存在实验证据。我们以前的理论认为,嵌入的比对图谱(以下简称“比对图谱”)提供了一种定量方法,能够将蛋白质的结构和功能属性以及它们的进化关系联系起来。比对简档的一个关键特征在于数据格式的互操作性(例如,比对信息、理化信息、基因组信息等)。事实上,我们已经证明了位置特定评分矩阵(PSSM)是通过定量测量嵌入的或未修改的序列比对来评分的信息性M维。此外,从这些比对中获得的信息是信息性的,即使在序列相似性(<25%的同一性)的“暮光区”中也是如此。虽然我们以前的嵌入策略是强大的,但它受到了(嵌入的和未修改的)污染比对和高计算成本的影响。在这里,我们描述了一种名为“自适应GDDA-BLAST”的启发式嵌入策略的逻辑和算法过程。平均而言,自适应GDDA-BLAST比我们以前的方法快19倍,但具有类似的灵敏度。此外,还提供了数据来证明嵌入比对测量在检测高度差异的蛋白质序列中的结构同源性和分离跨膜和锚蛋白重复结构域的二级结构元件方面的好处。总而言之,这些进步允许在足够大的数据集内进一步探索嵌入的比对数据空间,以最终得出相关的统计推断。我们表明,序列嵌入可以作为测量低同一性比对的工具之一,并将其并入基于PSSM的高性能比对配置文件中。
A major computational challenge in the genomic era is annotating structure/function to the vast quantities of sequence information that is now available. This problem is illustrated by the fact that most proteins lack comprehensive annotations, even when experimental evidence exists. We previously theorized that embedded-alignment profiles (simply “alignment profiles” hereafter) provide a quantitative method that is capable of relating the structural and functional properties of proteins, as well as their evolutionary relationships. A key feature of alignment profiles lies in the interoperability of data format (e.g., alignment information, physio-chemical information, genomic information, etc.). Indeed, we have demonstrated that the Position Specific Scoring Matrices (PSSMs) are an informative M-dimension that is scored by quantitatively measuring the embedded or unmodified sequence alignments. Moreover, the information obtained from these alignments is informative, and remains so even in the “twilight zone” of sequence similarity (<25% identity). Although our previous embedding strategy was powerful, it suffered from contaminating alignments (embedded AND unmodified) and high computational costs. Herein, we describe the logic and algorithmic process for a heuristic embedding strategy named “Adaptive GDDA-BLAST.” Adaptive GDDA-BLAST is, on average, up to 19 times faster than, but has similar sensitivity to our previous method. Further, data are provided to demonstrate the benefits of embedded-alignment measurements in terms of detecting structural homology in highly divergent protein sequences and isolating secondary structural elements of transmembrane and ankyrin-repeat domains. Together, these advances allow further exploration of the embedded alignment data space within sufficiently large data sets to eventually induce relevant statistical inferences. We show that sequence embedding could serve as one of the vehicles for measurement of low-identity alignments and for incorporation thereof into high-performance PSSM-based alignment profiles.
DOI: 10.1073/pnas.0803860105
发表时间: 2008-09-09
影响因子: 11.1
作者:
Chang, Gue Su;Hong, Yoojin;van Rossum, Damian B.
通讯作者: van Rossum, Damian B.
DOI: 10.1038/nature03340
发表时间: 2005-03-03
期刊: NATURE
影响因子: 64.8
作者:
van Rossum, DB;Patterson, RL;Snyder, SH
通讯作者: Snyder, SH
DOI: 10.1006/jmbi.2001.4495
发表时间: 2001-03-23
影响因子: 5.6
作者:
Blake, JD;Cohen, FE
通讯作者: Cohen, FE
DOI: 10.1006/jmbi.1998.2221
发表时间: 1998-12-11
影响因子: 5.6
作者:
Park, J;Karplus, K;Chothia, C
通讯作者: Chothia, C
DOI: 10.1002/prot.20830
发表时间: 2006-03-01
影响因子: 2.9
作者:
Kim, Y;Subramaniam, S
通讯作者: Subramaniam, S