课题基金 / 基金详情

项目摘要

项目成果

HUGH B NICHOLAS的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是许多研究子项目中的一个 由NIH/NCRR资助的中心赠款提供的资源。子项目和 研究者(PI)可能从另一个NIH来源获得了主要资金, 因此可以在其他CRISP条目中表示。所列机构为 研究中心,而研究中心不一定是研究者所在的机构。 在过去的一年里,我们开发了一种算法, 赋予相互作用必要特异性的生物大分子 在一个旁系同源分子家族的成员身上, 与不同的和独特的伙伴或基质一起或在其上起作用。 为 例如,每一种旁系同源的tRNA只选择性地与20种中的一种相互作用, 不同的氨酰tRNA合成酶。 每种旁系同源丝氨酸蛋白酶 仅涉及特定氨基酸的肽键,并且每种旁系同源物 异源三聚体G结合蛋白仅结合特异性受体并激活 只有特定的激酶或其他成员的特定信号通路。的 我们开发的算法识别了赋予这一点的序列特征, 对单个分子的特异性。 在应用此分析时,我们有 发现该分析不仅识别了赋予 所需的独特属性,但也确定了共同进化的合奏 我们在下面描述的序列元件。我们相信这些共同进化的 是生物学中的一个重要组成部分, 大分子微调其作用的特异性。 鉴定赋予作用特异性的序列元件, 生物大分子是通过将序列残基分成三个 第二类是我们研究的重点: 1. 结构所必需的高度保守序列残基, 活性的整个同源家庭的大分子。 2. 高度限制的序列残基,其维持了 旁系同源亚科的活动。 3. 可以自由变化的序列残基。 我们根据两个残基的数量将残基归为这三类 不同种类的熵与每个序列残基在一个多 序列比对 第一个是家庭相对熵, 在所有序列的比对中的特定位置处计算 在比对中(家族中的所有序列)。家庭相对熵 当比对中的所有序列都具有 在比对中的该位置处的相同种类的残基,并且该种类的残基 与其他可能的残留物相比是罕见的。 家庭相对熵是 计算为: 其中Pl是残基类型i在所述残基的特定位置中的分数, Q1是随机序列中预期的残基类型i的分数。 qi通常被视为适当的剩余类型的分数, 序列数据库 第二类熵是群交叉熵。 本集团 当只有一种剩余时,交叉熵达到最大值。 在组内发现和不同的单一种类的残留物被发现在 比对中的其余序列。 其计算为: sum(i){(qi-pi)*log2(pi/qi)} 其中Pl是残基类型i在所述残基的特定位置中的分数, 并且qi是预定义组中的序列的比对的分数, 在比对的特定位置上的残基类型I,对于不在 预定义的组。 这种形式的交叉熵是对称的,因此 可用作各种聚类过程中的距离度量。 第1类残基是那些具有高家族相对熵和 所有组的组交叉熵都很低。 第2类残基是那些 具有低的族相对熵和高的组交叉熵, 预定义组之一。 第3类残基是指 家庭相对熵和群体交叉熵较低。 我们一般 将高熵分数定义为3或更大的归一化Z分数,尽管 对于某些分析,低至2的值可能是有用的。 (The标准化Z分数 是原始分数减去平均分数,然后除以 分数的标准偏差)。 请注意,底层熵值为 不是正态分布,因此Z分数不应用于推断 统计显著性 该分析不特定于蛋白质或核酸序列。 这 让我能够验证这些方法是否适用于生物学上重要的 系统的答案已经知道从广泛的实验工作, 以及之前的分析。 我将新的分析应用于67的同一系统, 我之前用最初的简单计数模型分析过的tRNA (McClain和Nicholas,1987)。 实验工作证实了 McClain(1995)对这一早期分析进行了回顾。 新的, 基于信息的方法提供了相同的答案, 提高信噪比(Nicholas,1999)。 我把这个作为一个 在12月的“RNA信息的新兴来源”研讨会上发表演讲 在1998年。 改善的信噪比将使其更容易选择 通过实验室实验来测试答案。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. In the past year we have developed an algorithm that identifies the residues in biological macromolecules that confer the necessary specificity of interaction on the members of a paralogous family of molecules that carry out the same function with or upon different and distinct partners or substrates. For example, each paralogous tRNA interacts selectively with only one of twenty different aminoacyl tRNA synthetases. Each paralogous serine protease cleaves peptide bonds involving only particular amino acids, and each paralogous heterotrimeric G binding protein binds only a specific receptor and activates only particular kinases or other members of specific signaling pathways. The algorithm we have developed identifies the sequence features that confer this specificity on individual molecules. In applying this analysis, we have discovered that the analysis not only identifies sequence elements that confer the desired unique properties but also identifies ensembles of co-evolving sequence elements which we describe below. We believe these coevolving ensembles to be an important part of the mechanism by which biological macromolecules fine tune their specificity of action. The identification of sequence elements that confer specificity of action on biological macromolecules is achieved by dividing sequence residues into three categories, the second of which is the focus of our research: 1. Highly conserved sequence residues essential to the structure and activity of the entire homologous family of macromolecules. 2. Highly circumscribed sequence residues that maintain the specificity of the activity within the paralogous subfamilies. 3. Sequence residues that may vary freely. We assign residues to these three categories based on the amounts of two different kinds of entropy associated with each sequence residue in a multiple sequence alignment. The first is the family relative entropy, the entropy calculated at a particular position in the alignment over all of the sequences in the alignment (all the sequences in the family). The family relative entropy achieves its highest value when all of the sequences in an alignment have the same kind of residue at that position in the alignment and that kind of residue is rare compared to other possible residues. The family relative entropy is computed as: where pi is the fraction of residue type i in a particular position of the alignment and qi is the fraction of residue type i expected in random sequence. qi is usually taken as the fractions of residue types in an appropriate sequence database. The second kind of entropy considered is the group cross entropy. The group cross entropy achieves its highest value when only a single kind of residue is found within the group and a different single kind of residue is found in the rest of the sequences in the alignment. It is computed as: sum(i) {(qi-pi)*log2(pi/qi)} where pi is the fraction of residue type i in a particular position of the alignment for sequences in the predefined group and qi is the fraction of residue type i in a particular position of the alignment for sequences not in the predefined group. This form of the cross entropy is symmetric and hence usable as a distance measure in various clustering procedures. Category 1 residues are those that have a high family relative entropy and a low group cross entropy for all groups. Category 2 residues are those that have a low family relative entropy and a high group cross entropy for at least one of the predefined groups. Category 3 residues are those where both the family relative entropy and the group cross entropy are low. We generally define high entropy score to be a normalized Z score of 3 or greater, although for some analyses a value as low as 2 can be useful. (The normalized Z score is the raw score minus the average score and this difference divided by the standard deviation of the scores.) Note that the underlying entropy values are not normally distributed and thus the Z scores should not be used for inferring statistical significance. The analysis is not specific to either protein or nucleic acid sequences. This allowed me to check that the methods would work on a biologically important system where the answers were already known from extensive experimental work as well as previous analysis. I applied the new analysis to the same system of 67 tRNAs that I had analyzed earlier with the initial, simple counting models (McClain and Nicholas, 1987). The experimental work confirming the correctness of this earlier analysis is reviewed in McClain (1995). The new, information-based methods provided the same answers with a substantially improved signal to noise ratio (Nicholas, 1999). I presented this as an invited talk at the "Emerging Sources of RNA Information" workshop in December of 1998. The improved signal to noise ratio will make it easier select which answers to test by laboratory experiment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MARC: SUMMER INSTITUTE IN BIOINFORMATICS - JUNE 2010
  • 批准号:
    8364398
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
NRBSC/PSC PSC MARC INTERNS
  • 批准号:
    8364393
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2009
  • 批准号:
    8364394
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2008
  • 批准号:
    8171960
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2010
  • 负责人:
    HUGH B NICHOLAS
  • 依托单位:
海外基金