A Database Of Conserved Domain Alignments
A Database Of Conserved Domain Alignments
批准号:
6988471
负责人:
STEPHEN H. BRYANT
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
中文摘要
我们正在建立一个专家策划的蛋白质结构域比对数据库,描述蛋白质家族内的序列和3D结构保守性,并捕获这些家族内已知的功能多样性。
这些多重序列比对用于生成位置特异性得分矩阵(PSSM),其可以进而用于NCBI的基于网络的蛋白质分类资源。链接到保守域数据库(CDD)是默认从NCBI?的BLAST资源,http://www.ncbi.nlm.nih.gov/BLAST/,并从蛋白质记录在NCBI?的PubMed/http://www.ncbi.nlm.nih.gov/entrez/query.fcgi有关CDD和这些搜索服务的更多信息,请访问http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml。这些信息服务可用于鉴定蛋白质序列内的保守结构域。
策展领域模型总结了家族成员的已知功能,尽可能使用PubMed的相关引文。它们还通过序列和结构比对以及基于证据的相互作用位点特征提供位点特异性功能注释。CDD比对策展项目与早期的努力不同,它建立在两个基本方面:(i)尽可能以定量的方式使用3D结构信息来指导比对,以及(ii)家族和亚家族的明确层次结构,通过共同祖先的血统相关,反映每个领域的进化历史。当在结构域家族内已知3D结构时,该信息用于定义保守的3D核心结构,即必须在比对中包括的所有代表性序列中鉴定的一组无空位的块。使用线程或基于结构的比对算法将代表性序列与该核心结构进行比对,或者当已知多个结构时,通过结构-结构比对。这些程序确保了高比对准确性,如将注释准确转移到通过搜索识别的新家族成员所需。
明确的层次结构确定每个家庭的分子进化中的主要基因重复事件。我们的基本策略是使用域序列聚类方法与已知的域架构和同源性,以确定什么似乎是古老的正字法组。这些明确定义了总体“父”比对的注释“子”,并进而提供更具体的功能注释。CDD项目采用高水平的自动化,以产生基于结构的比对,识别候选的同源组,用新的序列和结构更新CDD比对,并将结果“发布”到Web服务器。这些算法和所需的相关软件在另一个项目“保守域数据库的对齐工具”LM 000045 -12中描述。这个项目描述了人类专家对CDD比对的管理。CDD策展人的作用是多方面的。他们首先必须调查相关的科学文献,以产生每个域家族的已知功能的简明摘要,并选择对NCBI用户有用的引文。的网络分类资源。策展人还必须检查自动序列和结构比较的结果,以推断保守核心块的位置,这是一个迭代过程,需要对消除不完整或错误的序列和结构数据进行判断。策展人还必须根据其他分子进化和聚类方法的一致结果,确定明显的同源组。到目前为止,策展人小组已经产生了大约1000个策划的CDD家族,这些家族现在可以通过NCBI的蛋白质分类服务器获得。
英文摘要
We are producing a database of expert-curated protein domain alignments, describing sequence and 3D-structure conservation within protein families and capturing the known functional diversity within these families.
These multiple sequence alignments are used to generate position-specific score matrices (PSSMs), that may in turn be used in NCBI's web-based protein classification resources. Links to the Conserved Domain Database (CDD) are made by default from NCBI?s BLAST resource, http://www.ncbi.nlm.nih.gov/BLAST/, and from protein records in NCBI?s PubMed/Entrez browser, http://www.ncbi.nlm.nih.gov/entrez/query.fcgi. Further information about CDD and these search services is available at http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml. These information services may be used to identify conserved domains within a protein sequence.
Curated domain models summarize the known functions of family members, using relevant citations from PubMed when possible. They also provide site-specific functional annotation, via sequence and structure alignments and via evidence-based interaction-site features. The CDD alignment curation project differs from earlier efforts, upon which it builds, in two fundamental ways: (i) 3D-structure information is used in a quantitative way, whenever possible, to guide alignments, and (ii) an explicit hierarchy of families and subfamilies, related by descent from a common ancestor, reflects the evolutionary history of each domain. When a 3D structure is known within a domain family, this information is used to define a conserved 3D core structure, a set of un-gapped blocks that must be identified in all representative sequences included in the alignment. Representative sequences are aligned to this core structure using threading or structure-based alignment algorithms or, when multiple structures are known, by structure-structure alignment. These procedures assure high alignment accuracy, as needed for accurate transfer of annotation to new family members identified by searching.
Explicit hierarchies identify major gene duplication events in the molecular evolution of each family. Our basic strategy is to use domain-sequence clustering methods together with known domain architecture and phylogeny to identify what appear to be ancient orthology groups. These define explicitly annotated "children" of the overall "parent" alignment, and in turn provide more specific functional annotation. The CDD project employs a high level of automation, to produce structure-based alignments, to identify candidate orthology groups, to update CDD alignments with new sequences and structures, and to "publish" the results to web servers. These algorithms and associated software required are described under another project, "Alignment Tools for the Conserved Domain Database," LM000045-12. This project describes human-expert curation of CDD alignments. The role of the CDD curators is multifaceted. They first of all must survey relevant scientific literature, to produce concise summaries of the known functions of each domain family and to choose citations useful to users of NCBI?s web-based classification resources. Curators must also examine the results of automated sequence and structure comparison to infer the location of conserved core blocks, an iterative process that requires judgment with respect to elimination of incomplete or erroneous sequence and structure data. Curators must also identify apparent orthology groups, based on the consensus of results from alternative molecular evolution and clustering methods. The curator group has so far produced about 1000 curated CDD families which are now available via NCBI's protein classification servers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
PROTEIN 3D STRUCTURE COMPARISON
-
批准号:6111066
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Internet Resources for Structural Bioinformatics
-
批准号:6554463
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
PROTEIN 3D STRUCTURE COMPARISON
-
批准号:6432751
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
INTERNET RESOURCES FOR MOLECULAR 3D STRUCTURE
-
批准号:6111065
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
INTERNET RESOURCES FOR MOLECULAR 3D STRUCTURE
-
批准号:6290484
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Internet Resources For Structural Bioinformatics
-
批准号:6843566
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Comparative Analysis of Protein 3-Dimensional Structure
-
批准号:6988454
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
INTERNET RESOURCES FOR MOLECULAR 3D STRUCTURE
-
批准号:6432750
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Comparative Analysis of Protein 3 Dimensional Structure
-
批准号:6554464
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
A Database Of Conserved Domain Alignments
-
批准号:6681403
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
A Database Of Conserved Domain Alignments
-
批准号:7316275
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Protien Threading Methods for a Conserved Domain Database
-
批准号:6432749
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
PROTEIN THREADING METHODS
-
批准号:6111064
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Comparative Analysis Of Protein 3-dimensional Structure
-
批准号:6843568
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Internet Resources For Structural Bioinformatics
-
批准号:7316234
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Bioinformatics Methods for Mass Spectra Analysis
-
批准号:7316288
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Comparative Analysis Of Protein 3-dimensional Structure
-
批准号:7316235
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
Alignment Methods For A Conserved Domain Database
-
批准号:6681331
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
A Database Of Conserved Domain Alignments
-
批准号:6843679
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
PROTEIN 3D STRUCTURE COMPARISON
-
批准号:6290485
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:STEPHEN H. BRYANT
-
依托单位:
海外基金