课题基金 / 基金详情

A Database Of Conserved Domain Alignments

A Database Of Conserved Domain Alignments
保守域比对数据库
批准号:
7316275
负责人:
STEPHEN H. BRYANT
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

STEPHEN H. BRYANT的其他基金

相似基金

相关文献

中文摘要
翻译
利用保守的结构域数据库(CDD)资源,我们正在建立一个专家精选的蛋白质结构域比对数据库。这种比对模型描述了蛋白质家族中的序列和3D结构保守,便于对保守的功能特征进行注释。比对模型还描述了领域家族中存在的可变性,有助于描述其功能多样性。 这个项目描述了人类专家对CDD比对的管理。CDD馆长的角色是多方面的。首先,他们必须查阅相关的科学文献,对每个领域家族的已知功能进行简明扼要的总结,研究现有的亚家族分类,并选择对用户有用的引文?S网络分类资源。馆长还必须检查自动序列和结构比较的结果,以推断保守的核心块的位置,这是一个迭代过程,需要关于消除不完整或错误的序列和结构数据的判断。策展人还必须根据可供选择的分子进化和聚类方法的结果的共识,确定明显的正畸类群。到目前为止,策展人小组已经制作了大约1500个经过策划的CDD家庭。精选和非精选的多重序列比对都用于生成特定位置的评分矩阵(PSSM),这些矩阵又可以用于NCBI的基于网络的蛋白质分类资源。 许多NCBI信息服务机构使用CDD来识别蛋白质序列中的保守结构域。例如,默认情况下,指向CDD的链接来自: 1)NCBI-S蛋白-BLAST资源,http://www.ncbi.nlm.nih.gov/BLAST/ 2)NCBI中的蛋白质?S Entrez Browser,http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Protein 3)在NCBI?S同源基因系统中的记录,http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=homologene. 欲了解有关CDD和这些搜索服务的更多信息,请访问http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml.。 经过策划的领域模型总结了家庭成员的已知功能,如果可能的话,使用PubMed的相关引用,并可能链接到NCBI书架上的资源以获取更多信息。它们还通过序列和结构比对以及预先记录的基于证据的特征,如相互作用或活性位点,提供特定于位点的功能注释。CDD比对整理项目不同于可比工作,它建立在两个基本方面:(I)尽可能以定量的方式使用3D结构信息来指导比对,以及(Ii)通过共同祖先的后代建立联系的明确的家族和子家族层次,反映每个领域超级家族的进化历史。 当结构域家族内已知至少一个3D结构时,该信息用于定义保守的同源核心结构,即必须在包括在比对中的所有代表性序列中鉴定的一组无间隙的块。使用结构信息比对算法将代表性序列比对到该核心结构,或者当已知多个3D结构时,使用从结构叠加获得的比对。这些步骤可确保较高的对齐精度,以将注释准确传递给通过搜索标识的新族成员。代表性的序列是从一组首选的分类节点中挑选出来的,因此结构域排列代表了一个家族的分类跨度,这反过来又表明了它的明显进化年龄。 明确的层次结构确定了每个家族分子进化中的主要基因复制事件。我们的基本策略是使用域序列聚类方法,结合已知的域体系结构和系统发育来识别看起来像是古代正字法组的东西。它们定义了整个“父”对齐的显式注释“子”,进而提供了更具体的功能注释。CDD项目采用高水平的自动化,以产生基于结构的比对,识别候选的正畸分组,用新的序列和结构更新CDD比对,并将结果发布到网络服务器。这些算法和所需的相关软件在另一个项目“保守域数据库的比对方法”中进行了描述。
英文摘要
With the Conserved Domain Database (CDD) resource we are producing a database of expert-curated protein domain alignments. Such alignment models describe the sequence and 3D-structure conservation within protein families, facilitating the annotation of conserved functional features. The alignment models also describe the variability present in a domain family, facilitating the depiction of its functional diversity. This project describes curation of CDD alignments by human experts. The role of the CDD curators is multifaceted. First of all they must survey relevant scientific literature, to produce concise summaries of the known functions of each domain family, to study existing sub-family classifications, and to choose citations useful to users of NCBI?s web-based classification resources. Curators must also examine the results of automated sequence and structure comparison to infer the location of conserved core blocks, an iterative process that requires judgment with respect to elimination of incomplete or erroneous sequence and structure data. Curators must also identify apparent orthology groups, based on the consensus of results from alternative molecular evolution and clustering methods. The curator group has so far produced about 1500 curated CDD families. Both curated and un-curated multiple sequence alignments are used to generate position-specific scoring matrices (PSSMs), which may in turn be used in NCBI's web-based protein classification resources. A number of NCBI information services use CDD to identify conserved domains within protein sequences. Links to CDD are made, for example, by default from: 1) NCBI?s protein-BLAST resource, http://www.ncbi.nlm.nih.gov/BLAST/ 2) proteins in NCBI?s Entrez browser, http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Protein 3) records in NCBI?s HomoloGene system, http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=homologene. Further information about CDD and these search services is available at http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml. Curated domain models summarize the known functions of family members, using relevant citations from PubMed when possible, and may link to resources on the NCBI Bookshelf for further information. They also provide site-specific functional annotation, via sequence and structure alignments and via pre-recorded evidence-based features, such as interaction or active sites. The CDD alignment curation project differs from comparable efforts, upon which it builds, in two fundamental ways: (i) 3D-structure information is used in a quantitative way, whenever possible, to guide the alignments, and (ii) an explicit hierarchy of families and subfamilies, related by descend from a common ancestor, reflects the evolutionary history of each domain super-family. When at least one 3D structure is known within a domain family, this information is used to define the conserved homologous core structure, a set of un-gapped blocks that must be identified in all representative sequences included in the alignment. Representative sequences are aligned to this core structure using structure-informed alignment algorithms or, when multiple 3D structures are known, alignments obtained from structure superposition. These procedures assure high alignment accuracy, as needed for accurate transfer of annotation to new family members identified by searching. Representative sequences are picked from a set of ?preferred taxonomy nodes?, so that the domain alignments represent the taxonomic span of a family, which in turn indicates its apparent evolutionary age. Explicit hierarchies identify major gene duplication events in the molecular evolution of each family. Our basic strategy is to use domain-sequence clustering methods together with known domain architecture and phylogeny to identify what appear to be ancient orthology groups. These define explicitly annotated "children" of the overall "parent" alignment, and in turn provide more specific functional annotation. The CDD project employs a high level of automation, to produce structure-based alignments, to identify candidate orthology groups, to update CDD alignments with new sequences and structures, and to "publish" the results to web servers. These algorithms and associated software required are described under another project, "Alignment methods for a conserved domain database".
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
PROTEIN 3D STRUCTURE COMPARISON
  • 批准号:
    6111066
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN H. BRYANT
  • 依托单位:
Internet Resources for Structural Bioinformatics
  • 批准号:
    6554463
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN H. BRYANT
  • 依托单位:
PROTEIN 3D STRUCTURE COMPARISON
  • 批准号:
    6432751
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN H. BRYANT
  • 依托单位:
INTERNET RESOURCES FOR MOLECULAR 3D STRUCTURE
  • 批准号:
    6111065
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN H. BRYANT
  • 依托单位:
海外基金