课题基金 / 基金详情

项目摘要

项目成果

HUGH B NICHOLAS的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是许多研究子项目中利用 资源由NIH/NCRR资助的中心拨款提供。子项目和 调查员(PI)可能从NIH的另一个来源获得了主要资金, 并因此可以在其他清晰的条目中表示。列出的机构是 该中心不一定是调查人员的机构。 在过去的一年里,我们开发了一种算法,可以识别出 赋予相互作用必要的专一性的生物大分子 在执行相同任务的平行分子家族的成员上 与不同的、不同的伙伴或底物一起或在其上发挥作用。为 例如,每个类似的tRNA只与20个中的一个有选择地相互作用 不同的氨酰tRNA合成酶。每个平行的丝氨酸蛋白酶都会裂解 只包含特定氨基酸的多肽键,并且每个氨基酸都是平行的 异三聚体G结合蛋白仅与特定受体结合并激活 只有特定的激酶或特定信号通路的其他成员。这个 我们开发的算法识别了实现这一点的序列特征 对单个分子的专一性。在应用这一分析时,我们有 发现该分析不仅识别了赋予 所需的独特特性,但也确定了共同进化的系综 我们在下面描述的序列元素。我们相信这些共同进化 集合是生物学机制的重要组成部分 大分子微调其作用的专一性。 识别赋予作用的特异性的序列元素 生物大分子是通过将序列残基分成三部分来获得的 类别,其中第二个是我们研究的重点: 1.高度保守的序列残基,对结构和 整个同源大分子家族的活性。 2.高度受限的序列残基,保持其特异性 在类似的亚科中的活动。 3.可自由变化的序列残基。 我们将残基分配给这三个类别,根据两个 与多个序列中的每个序列残留物相关联的不同种类的熵 序列比对。第一个是家族的相对熵,即 在所有序列的比对中的特定位置计算 在比对中(家族中的所有序列)。家族相对熵 当比对中的所有序列都具有 在对齐位置上的相同类型的残基和那种残基 与其他可能的残留物相比是罕见的。家族的相对熵是 计算方式为: 其中pi是第I类残基在特定位置的分数 比对和齐是随机序列中残基类型I预期的分数。 气通常作为残基类型的分数在适当的 序列数据库。 考虑的第二种熵是群交叉熵。这群人 当只有一种残差时,交叉熵达到最大值 在该群中发现了另一种不同种类的残留物 比对中的其余序列。它的计算公式为: Sum(I){(qi-pi)*log2(pi/qi)} 其中pi是第I类残基在特定位置的分数 预定义的组和QI中的序列的比对是 在序列比对的特定位置上的I型残基 预定义的组。这种形式的交叉熵是对称的,因此 在各种聚类过程中可用作距离度量。 1类残基是那些具有较高家族相对熵和 所有组的组交叉熵都很低。第2类残留物是指那些 具有低的家族相对熵和高的群交叉熵至少 预定义的组之一。第3类残留物是指既有 家族相对熵和群交叉熵较低。我们一般 将高熵分数定义为3或更大的归一化Z分数,尽管 对于某些分析,低至2的值可能很有用。(归一化Z分数 是原始分数减去平均分数,这个差除以 分数的标准差。)请注意,基本的熵值为 不是正态分布的,因此不应使用Z分数进行推断 统计学意义。 这种分析既不针对蛋白质序列,也不针对核酸序列。这 允许我检查这些方法是否适用于一种具有重要生物学意义的 从广泛的实验工作中已经知道答案的系统,如 和之前的分析一样。我将新的分析应用于67个相同的系统 我早些时候用初始的简单计数模型分析过的tRNA (麦克莱恩和尼古拉斯,1987)。验证其正确性的实验工作 在McClain(1995)中对这一早期的分析进行了回顾。新的, 基于信息的方法提供了相同的答案,基本上 改善了信噪比(Nicholas,1999)。我把这个作为一种 12月在“RNA信息的新兴来源”研讨会上的特邀演讲 1998年。改善的信噪比将使其更容易选择 通过实验室实验对测试的答案。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. In the past year we have developed an algorithm that identifies the residues in biological macromolecules that confer the necessary specificity of interaction on the members of a paralogous family of molecules that carry out the same function with or upon different and distinct partners or substrates. For example, each paralogous tRNA interacts selectively with only one of twenty different aminoacyl tRNA synthetases. Each paralogous serine protease cleaves peptide bonds involving only particular amino acids, and each paralogous heterotrimeric G binding protein binds only a specific receptor and activates only particular kinases or other members of specific signaling pathways. The algorithm we have developed identifies the sequence features that confer this specificity on individual molecules. In applying this analysis, we have discovered that the analysis not only identifies sequence elements that confer the desired unique properties but also identifies ensembles of co-evolving sequence elements which we describe below. We believe these coevolving ensembles to be an important part of the mechanism by which biological macromolecules fine tune their specificity of action. The identification of sequence elements that confer specificity of action on biological macromolecules is achieved by dividing sequence residues into three categories, the second of which is the focus of our research: 1. Highly conserved sequence residues essential to the structure and activity of the entire homologous family of macromolecules. 2. Highly circumscribed sequence residues that maintain the specificity of the activity within the paralogous subfamilies. 3. Sequence residues that may vary freely. We assign residues to these three categories based on the amounts of two different kinds of entropy associated with each sequence residue in a multiple sequence alignment. The first is the family relative entropy, the entropy calculated at a particular position in the alignment over all of the sequences in the alignment (all the sequences in the family). The family relative entropy achieves its highest value when all of the sequences in an alignment have the same kind of residue at that position in the alignment and that kind of residue is rare compared to other possible residues. The family relative entropy is computed as: where pi is the fraction of residue type i in a particular position of the alignment and qi is the fraction of residue type i expected in random sequence. qi is usually taken as the fractions of residue types in an appropriate sequence database. The second kind of entropy considered is the group cross entropy. The group cross entropy achieves its highest value when only a single kind of residue is found within the group and a different single kind of residue is found in the rest of the sequences in the alignment. It is computed as: sum(i) {(qi-pi)*log2(pi/qi)} where pi is the fraction of residue type i in a particular position of the alignment for sequences in the predefined group and qi is the fraction of residue type i in a particular position of the alignment for sequences not in the predefined group. This form of the cross entropy is symmetric and hence usable as a distance measure in various clustering procedures. Category 1 residues are those that have a high family relative entropy and a low group cross entropy for all groups. Category 2 residues are those that have a low family relative entropy and a high group cross entropy for at least one of the predefined groups. Category 3 residues are those where both the family relative entropy and the group cross entropy are low. We generally define high entropy score to be a normalized Z score of 3 or greater, although for some analyses a value as low as 2 can be useful. (The normalized Z score is the raw score minus the average score and this difference divided by the standard deviation of the scores.) Note that the underlying entropy values are not normally distributed and thus the Z scores should not be used for inferring statistical significance. The analysis is not specific to either protein or nucleic acid sequences. This allowed me to check that the methods would work on a biologically important system where the answers were already known from extensive experimental work as well as previous analysis. I applied the new analysis to the same system of 67 tRNAs that I had analyzed earlier with the initial, simple counting models (McClain and Nicholas, 1987). The experimental work confirming the correctness of this earlier analysis is reviewed in McClain (1995). The new, information-based methods provided the same answers with a substantially improved signal to noise ratio (Nicholas, 1999). I presented this as an invited talk at the "Emerging Sources of RNA Information" workshop in December of 1998. The improved signal to noise ratio will make it easier select which answers to test by laboratory experiment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MARC: SUMMER INSTITUTE IN BIOINFORMATICS - JUNE 2010
  • 批准号:
    8364398
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
NRBSC/PSC PSC MARC INTERNS
  • 批准号:
    8364393
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2009
  • 批准号:
    8364394
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2008
  • 批准号:
    8171960
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2010
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
海外基金