The CaspBase: a curated database for evolutionary biochemical studies of caspase functional divergence and ancestral sequence inference

The CaspBase: a curated database for evolutionary biochemical studies of caspase functional divergence and ancestral sequence inference
复制标题

DOI:
10.1002/pro.3494
复制
发表时间:
2018-10-01
期刊:
影响因子:
8
通讯作者:
Clark, A. Clay
Clark, A. Clay
中科院分区:
生物学3区
文献类型:
--
作者:
Grinshpon, Robert D.;Williford, Anna;Clark, A. Clay

文献摘要

被引文献

相似文献

序列数据库是当代科学家的有力工具。然而,公共数据库中的大多数功能注释都是通过计算确定的,而不是由人类专家验证的。虽然从计算研究中产生的假设现在可以进行实验,但结果的质量取决于输入数据的质量。我们开发了CaspBase,以加快注释的caspase序列的高质量数据集编译,最大限度地提高系统发育信号,并减少公共数据库的噪音。我们描述了我们的方法,策展的CaspBase和研究人员如何可以获得序列。我们开发CaspBase的直接目标是优化Caspases的祖先蛋白重建(APR),我们证明了CaspBase在APR研究中的实用性。我们还开发了共同的位置(CP)系统比较人类胱天蛋白酶家族旁系同源物,并建议CP系统作为更新目前的报告方法的胱天蛋白酶的氨基酸位置。我们提出了一个标准化的多序列比对(MSA)的CP系统,并显示使用大型数据库,如CaspBase在定义蛋白质中的结构位置的优势。虽然这里描述的结果属于半胱天冬酶的进化和结构功能的研究,该方法可以适用于任何基因家族。
Sequence databases are powerful tools for the contemporary scientists' toolkit. However, most functional annotations in public databases are determined computationally and are not verified by a human expert. While hypotheses generated from computational studies are now amenable to experimentation, the quality of the results relies on the quality of input data. We developed the CaspBase to expedite high-quality dataset compilation of annotated caspase sequences, to maximize phylogenetic signal, and to reduce the noise contributed from public databanks. We describe our methods of curation for the CaspBase and how researchers can acquire sequences from . Our immediate goal for developing the CaspBase was to optimize the ancestral protein reconstruction (APR) of caspases, and we demonstrate the utility of the CaspBase in APR studies. We also developed the Common Position (CP) system for comparing human caspase family paralogs and suggest the CP system as an update to current reporting methods of caspase amino acid positions. We present a standardized multiple sequence alignment (MSA) for the CP system and show the advantage of using large databases such as the CaspBase in defining structural positions in proteins. Although the results described here pertain to caspase evolution and structure-function studies, the methods can be adapted to any gene family.