CORE RESEARCH GRANT
CORE RESEARCH GRANT
批准号:
7723100
负责人:
HUGH B NICHOLAS
金额:
$0.05万
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-08-01 至 2009-07-31
关键词:
AlgorithmsAmino AcidsAmino Acyl-tRNA SynthetasesBase SequenceBinding ProteinsBiologicalCategoriesCleaved cellComputer Retrieval of Information on Scientific Projects DatabaseConserved SequenceCountDatabasesEducational workshopElementsEntropyFamilyFundingGrantIndividualInstitutionLaboratoriesMeasuresMethodsModelingNoiseNucleic acid sequencingObject AttachmentPeptidesPhosphotransferasesPositioning AttributeProceduresPropertyProtein BindingProteinsQiRNARelative (related person)ResearchResearch PersonnelResearch Project GrantsResourcesRestScoreSequence AlignmentSerine ProteaseSignal PathwaySignal TransductionSourceSpecificityStandards of Weights and MeasuresStructureSumSystemTestingTransfer RNAUnited States National Institutes of HealthWorkbasedesireimprovedmacromoleculememberreceptorresearch study
中文摘要
这个子项目是许多研究子项目中的一个
由NIH/NCRR资助的中心赠款提供的资源。子项目和
研究者(PI)可能从另一个NIH来源获得了主要资金,
因此可以在其他CRISP条目中表示。所列机构为
研究中心,而研究中心不一定是研究者所在的机构。
在过去的一年里,我们开发了一种算法,
赋予相互作用必要特异性的生物大分子
在一个旁系同源分子家族的成员身上,
与不同的和独特的伙伴或基质一起或在其上起作用。 为
例如,每一种旁系同源的tRNA只选择性地与20种中的一种相互作用,
不同的氨酰tRNA合成酶。 每种旁系同源丝氨酸蛋白酶
仅涉及特定氨基酸的肽键,并且每种旁系同源物
异源三聚体G结合蛋白仅结合特异性受体并激活
只有特定的激酶或其他成员的特定信号通路。的
我们开发的算法识别了赋予这一点的序列特征,
对单个分子的特异性。 在应用此分析时,我们有
发现该分析不仅识别了赋予
所需的独特属性,但也确定了共同进化的合奏
我们在下面描述的序列元件。我们相信这些共同进化的
是生物学中的一个重要组成部分,
大分子微调其作用的特异性。
鉴定赋予作用特异性的序列元件,
生物大分子是通过将序列残基分成三个
第二类是我们研究的重点:
1. 结构所必需的高度保守序列残基,
活性的整个同源家庭的大分子。
2. 高度限制的序列残基,其维持了
旁系同源亚科的活动。
3. 可以自由变化的序列残基。
我们根据两个残基的数量将残基归为这三类
不同种类的熵与每个序列残基在一个多
序列比对 第一个是家庭相对熵,
在所有序列的比对中的特定位置处计算
在比对中(家族中的所有序列)。家庭相对熵
当比对中的所有序列都具有
在比对中的该位置处的相同种类的残基,并且该种类的残基
与其他可能的残留物相比是罕见的。 家庭相对熵是
计算为:
其中Pl是残基类型i在所述残基的特定位置中的分数,
Q1是随机序列中预期的残基类型i的分数。
qi通常被视为适当的剩余类型的分数,
序列数据库
第二类熵是群交叉熵。 本集团
当只有一种剩余时,交叉熵达到最大值。
在组内发现和不同的单一种类的残留物被发现在
比对中的其余序列。 其计算为:
sum(i){(qi-pi)*log2(pi/qi)}
其中Pl是残基类型i在所述残基的特定位置中的分数,
并且qi是预定义组中的序列的比对的分数,
在比对的特定位置上的残基类型I,对于不在
预定义的组。 这种形式的交叉熵是对称的,因此
可用作各种聚类过程中的距离度量。
第1类残基是那些具有高家族相对熵和
所有组的组交叉熵都很低。 第2类残基是那些
具有低的族相对熵和高的组交叉熵,
预定义组之一。 第3类残基是指
家庭相对熵和群体交叉熵较低。 我们一般
将高熵分数定义为3或更大的归一化Z分数,尽管
对于某些分析,低至2的值可能是有用的。 (The标准化Z分数
是原始分数减去平均分数,然后除以
分数的标准偏差)。 请注意,底层熵值为
不是正态分布,因此Z分数不应用于推断
统计显著性
该分析不特定于蛋白质或核酸序列。 这
让我能够验证这些方法是否适用于生物学上重要的
系统的答案已经知道从广泛的实验工作,
以及之前的分析。 我将新的分析应用于67的同一系统,
我之前用最初的简单计数模型分析过的tRNA
(McClain和Nicholas,1987)。 实验工作证实了
McClain(1995)对这一早期分析进行了回顾。 新的,
基于信息的方法提供了相同的答案,
提高信噪比(Nicholas,1999)。 我把这个作为一个
在12月的“RNA信息的新兴来源”研讨会上发表演讲
在1998年。 改善的信噪比将使其更容易选择
通过实验室实验来测试答案。
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
In the past year we have developed an algorithm that identifies the residues in
biological macromolecules that confer the necessary specificity of interaction
on the members of a paralogous family of molecules that carry out the same
function with or upon different and distinct partners or substrates. For
example, each paralogous tRNA interacts selectively with only one of twenty
different aminoacyl tRNA synthetases. Each paralogous serine protease cleaves
peptide bonds involving only particular amino acids, and each paralogous
heterotrimeric G binding protein binds only a specific receptor and activates
only particular kinases or other members of specific signaling pathways. The
algorithm we have developed identifies the sequence features that confer this
specificity on individual molecules. In applying this analysis, we have
discovered that the analysis not only identifies sequence elements that confer
the desired unique properties but also identifies ensembles of co-evolving
sequence elements which we describe below. We believe these coevolving
ensembles to be an important part of the mechanism by which biological
macromolecules fine tune their specificity of action.
The identification of sequence elements that confer specificity of action on
biological macromolecules is achieved by dividing sequence residues into three
categories, the second of which is the focus of our research:
1. Highly conserved sequence residues essential to the structure and
activity of the entire homologous family of macromolecules.
2. Highly circumscribed sequence residues that maintain the specificity of
the activity within the paralogous subfamilies.
3. Sequence residues that may vary freely.
We assign residues to these three categories based on the amounts of two
different kinds of entropy associated with each sequence residue in a multiple
sequence alignment. The first is the family relative entropy, the entropy
calculated at a particular position in the alignment over all of the sequences
in the alignment (all the sequences in the family). The family relative entropy
achieves its highest value when all of the sequences in an alignment have the
same kind of residue at that position in the alignment and that kind of residue
is rare compared to other possible residues. The family relative entropy is
computed as:
where pi is the fraction of residue type i in a particular position of the
alignment and qi is the fraction of residue type i expected in random sequence.
qi is usually taken as the fractions of residue types in an appropriate
sequence database.
The second kind of entropy considered is the group cross entropy. The group
cross entropy achieves its highest value when only a single kind of residue is
found within the group and a different single kind of residue is found in the
rest of the sequences in the alignment. It is computed as:
sum(i) {(qi-pi)*log2(pi/qi)}
where pi is the fraction of residue type i in a particular position of the
alignment for sequences in the predefined group and qi is the fraction of
residue type i in a particular position of the alignment for sequences not in
the predefined group. This form of the cross entropy is symmetric and hence
usable as a distance measure in various clustering procedures.
Category 1 residues are those that have a high family relative entropy and a
low group cross entropy for all groups. Category 2 residues are those that
have a low family relative entropy and a high group cross entropy for at least
one of the predefined groups. Category 3 residues are those where both the
family relative entropy and the group cross entropy are low. We generally
define high entropy score to be a normalized Z score of 3 or greater, although
for some analyses a value as low as 2 can be useful. (The normalized Z score
is the raw score minus the average score and this difference divided by the
standard deviation of the scores.) Note that the underlying entropy values are
not normally distributed and thus the Z scores should not be used for inferring
statistical significance.
The analysis is not specific to either protein or nucleic acid sequences. This
allowed me to check that the methods would work on a biologically important
system where the answers were already known from extensive experimental work as
well as previous analysis. I applied the new analysis to the same system of 67
tRNAs that I had analyzed earlier with the initial, simple counting models
(McClain and Nicholas, 1987). The experimental work confirming the correctness
of this earlier analysis is reviewed in McClain (1995). The new,
information-based methods provided the same answers with a substantially
improved signal to noise ratio (Nicholas, 1999). I presented this as an
invited talk at the "Emerging Sources of RNA Information" workshop in December
of 1998. The improved signal to noise ratio will make it easier select which
answers to test by laboratory experiment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MARC: SUMMER INSTITUTE IN BIOINFORMATICS - JUNE 2010
-
批准号:8364398
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2011
-
负责人:HUGH B NICHOLAS
-
依托单位:
NRBSC/PSC PSC MARC INTERNS
-
批准号:8364393
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2011
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2009
-
批准号:8364394
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2011
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2008
-
批准号:8171960
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2010
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2009
-
批准号:8171961
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2010
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC: SUMMER INSTITUTE IN BIOINFORMATICS - JUNE 2010
-
批准号:8171965
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2010
-
负责人:HUGH B NICHOLAS
-
依托单位:
NRBSC/PSC PSC MARC INTERNS
-
批准号:8171958
-
项目类别:
-
资助金额:$0.11万
-
财政年份:2010
-
负责人:HUGH B NICHOLAS
-
依托单位:
PEURTO RICO BIOINFORMATICS WORKSHOP (APRIL 27-MAY1 2009)
-
批准号:7956382
-
项目类别:
-
资助金额:$8.71万
-
财政年份:2009
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 2008
-
批准号:7956381
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:HUGH B NICHOLAS
-
依托单位:
JACKSON STATE BIOINFORMATICS WORKSHOP (APRIL 16-17, 2007)
-
批准号:7956169
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 9 - 20, 2007
-
批准号:7956289
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:HUGH B NICHOLAS
-
依托单位:
ANALYSIS OF TYROSINE SULFATION: HIV
-
批准号:7956091
-
项目类别:
-
资助金额:$0.08万
-
财政年份:2009
-
负责人:HUGH B NICHOLAS
-
依托单位:
ANALYSIS OF TYROSINE SULFATION
-
批准号:7723138
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 9 - 20, 2007
-
批准号:7723430
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:HUGH B NICHOLAS
-
依托单位:
PARALLELIZING COMPUTATIONALLY INTENSIVE PHYLOGENETIC ANALYSIS ROUTINES
-
批准号:7723197
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:HUGH B NICHOLAS
-
依托单位:
MARC - DEVELOPING BIOINFORMATICS PROGRAMS, JULY 17-28, 2006
-
批准号:7723305
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:HUGH B NICHOLAS
-
依托单位:
JACKSON STATE BIOINFORMATICS WORKSHOP (APRIL 16-17, 2007)
-
批准号:7723307
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:HUGH B NICHOLAS
-
依托单位:
CORE RESEARCH GRANT
-
批准号:7601261
-
项目类别:
-
资助金额:$0.04万
-
财政年份:2007
-
负责人:HUGH B NICHOLAS
-
依托单位:
PREDICTING POTENTIAL OF VACCINE DEVELOPMENT AGAINST LEISHMANIA (L) MEXICANA PAR
-
批准号:7601411
-
项目类别:
-
资助金额:$0.03万
-
财政年份:2007
-
负责人:HUGH B NICHOLAS
-
依托单位:
PARALLELIZING COMPUTATIONALLY INTENSIVE PHYLOGENETIC ANALYSIS ROUTINES
-
批准号:7601450
-
项目类别:
-
资助金额:$0.03万
-
财政年份:2007
-
负责人:HUGH B NICHOLAS
-
依托单位:
海外基金