课题基金 / 基金详情

项目摘要

项目成果

PATRICIA CLEMENT BABBITT的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是利用资源的许多研究子项目之一。 由NIH/NCRR资助的中心拨款提供。对子项目的主要支持 子项目的首席调查员可能是由其他来源提供的, 包括美国国立卫生研究院的其他来源。为子项目列出的总成本可能 表示该子项目使用的中心基础设施的估计数量, 不是由NCRR赠款提供给次级项目或次级项目工作人员的直接资金。 结构-功能联动利用的一个重大悬而未决的问题 计算预测是,虽然我们可以准确地将蛋白质聚集在一起 具有良好统计意义的序列和结构 相似性度量的类型,这些集群如何链接到功能类 目前还不清楚。尽管像直系同源预测这样的简单方法可以 对相似或相似的序列获得良好的结果 包含区分功能类的容易识别的主题, 对于许多蛋白质超家族来说,成功的预测远非微不足道。 这就是SFLD中功能多样化的超级家庭的情况。 这些是同源的一组酶,它们执行不同的化学 转换,使用不同的底物,但都共享特定的 化学功能或部分反应。SFLD的主要用途 是为了帮助研究人员管理这些类型的超级家庭, 帮助识别这些超级家庭的新成员,并 为这些酶提供明确的结构-功能图谱。因为 一个给定的超级家族中的不同功能家族看起来很相似,但 执行不同的具体反应,它们很难注释和 很容易被误注解,在 档案数据库:Genbank、NR和TREMBL。因为序列信息是 仍有大量可用,需要自动化方法来 用新确定的序列更新SFLD超家族并分配 把他们送到适当的功能家庭。显然,改进的方法 迫切需要实现这些职能分配。 制定一种实现这一目标的方法一直是 RBVI与觉醒的Jacquelyn Fetrow教授小组合作 森林大学。主动部位分析方法由Dr。 Fetrow现在已经与巴比特人开发的方法集成在一起 实验室,遗传算法搜索结构中的模式:喘气,TO 自动确定能够区分新的 用于自动分配序列的超家族成员 到他们所属的特定功能家庭。喘息将会是 结合Fetrow的方法创建序列和结构主题 SFLD数据的自动集群。该方法的核心要素包括 一种称为“模糊函数形式”的模体生成技术(FFF), 由蛋白质活性位点结构搜索(PASSS)工具实现,以及 Deacon Active Site Profiler(DASP),它使用三维或 基于结构的活性部位分析,以确定位于 活动场地周围的空间环境。PASSS使用FFF 技术,通过蛋白质之间的距离来描述蛋白质的功能位点 对功能部位重要的三个关键残基的α碳 化学和相邻残基的阿尔法碳。基于 功能相关蛋白质应该具有结构的前提 在功能位点上的相似性,PASSS将相关蛋白质返回到 正在启动已知功能站点。DASP对此进行了扩展,提取了 在每个关键残基附近发现的残基 蛋白质,从这些片段中创建模体,并使用这些片段 搜索数据库中的所有序列以返回可能共享的蛋白质 此函数。以迭代的方式一起使用这些工具, 提供了一种快速的方法,可以从功能上对两者进行表征 结构和序列。 该项目的初步结果显示,在 区分烯醇化酶和激酶功能不同的家族 超级大家庭。前者是SFLD中注解的超家族之一 对于这种类型的自动化测试系统来说,这是一个具有挑战性的测试系统 努力。这条管道现在正在应用于Kinase超家族,以努力使用自动化方法将这个超家族添加到SFLD中。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. Primary support for the subproject and the subproject's principal investigator may have been provided by other sources, including other NIH sources. The Total Cost listed for the subproject likely represents the estimated amount of Center infrastructure utilized by the subproject, not direct funding provided by the NCRR grant to the subproject or subproject staff. A major unsolved problem for structure-function linkage using computational prediction is that while we can accurately cluster protein sequences and structures with good statistical significance based on many types of similarity metrics, how those clusters link to functional classes is not clear. Although simple approaches such as ortholog prediction can achieve good results for sequences that are closely similar or that contain readily identifiable motifs that distinguish functional classes, for many protein superfamilies successful prediction is far from trivial. This is the case for the functionally diverse superfamilies in the SFLD. These are homologous sets of enzymes that carry out different chemical transformations, using different substrates, but all share a specific chemical functionality or partial reaction. The main purpose of the SFLD is to aid researchers in the curation of these types of superfamilies, to help in the identification of new members of these superfamilies, and to provide an explicit structure-function mapping for these enzymes. Because the different functional families in a given superfamily look similar but perform different specific reactions, they are difficult to annotate and easy to misannotate, showing levels of misannotation as high as 80% in the archival databases Genbank NR and TrEMBL. Because sequence information is still coming available in large volumes, automated methods are required to update the SFLD superfamilies with newly determined sequences and assign them to the appropriate functional families. Clearly, improved methods for achieving these functional assignments are urgently needed. Development of an approach to achieve this has been a major focus of the RBVI in collaboration with the group of Prof. Jacquelyn Fetrow of Wake Forest University. The active site profiling methods developed by Dr. Fetrow have now been integrated with an approach developed in the Babbitt lab, Genetic Algorithm Search for Patterns in Structures: GASPS, to automatically determine 3D templates capable of distinguishing new superfamily members for the purpose of automatically assigning sequences to the specific functional families to which they belong. GASPS will be combined with Fetrow's methods to create sequence and structural motifs for automated clustering of SFLD data. The core elements of the method include a motif-generating technology called "Fuzzy Functional Forms", (FFF), implemented by the tool Protein Active Site Structure Search (PASSS), and the Deacon Active Site Profiler (DASP) which uses three-dimensional, or structure-based, active-site profiling to identify residues located in the spatial environment around the active site. PASSS uses the FFF technology, describing a proteins functional site by the distances between the alpha carbons of three key residues important to the functional site chemistry and the alpha carbons of adjacent residues. Based on the premise that functionally related proteins should have structural similarity at the functional site, PASSS returns related proteins to the starting known functional site. DASP expands on this, extracting the residues that are found in the vicinity of the key residues for each protein, creating motifs from these fragments, and using these fragments to search all sequences in a database to return proteins that may share this function. Use of these tools together, and in an iterative fashion, provides a quick method to putatively functionally characterize both structures and sequences. Preliminary results from this project show exceptional accuracy in distinguishing functionally diverse families in the enolase and the kinase superfamily. The former is one of the annotated superfamilies in the SFLD that serves as a challenging test system for this type of automated effort. This pipeline is now being applied to the Kinase superfamily in an effort to add this superfamily to the SFLD using an automated approach.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
THE STRUCTURE-FUNCTION LINKAGE DATABASE
LAYING THE FOUNDATIONS FOR GENOMIC ENZYMOLOGY
ACTIVE SITE SIGNATURES FOR SFLD: ENOLASE SUPERFAMILY
ACTIVE SITE SIGNATURES FOR AUTOMATIC UPDATES OF SFLD SUPERFAMILIES
海外基金