COMPUTER ANALYSIS OF LOW-COMPLEXITY AMINO ACID AND NUCLEOTIDE SEQUENCES
COMPUTER ANALYSIS OF LOW-COMPLEXITY AMINO ACID AND NUCLEOTIDE SEQUENCES
批准号:
6162792
负责人:
J C WOOTTON
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
中文摘要
该项目的目标是定义、分类和
使用计算分析来分析蛋白质片段
和显示组成偏差的核苷酸序列或
令人难以置信的低构图复杂性。蛋白质中
序列,其中包括丰富的残基簇
主要是一种或几种氨基酸,通常
含有这些非周期性的均聚束或马赛克
低周期重复的模式和节段。其他常见
例子包括长的非肾小球结构域。这个
氨基酸和氨基酸中偏向片段的丰度
核苷酸序列数据库已经确定,并且
它们的性质与证据有关
生物功能。
局部成分的不同形式定义
复杂性被用来无偏见地识别
低复杂性分段,在不同的严格程度上。
算法经过改进以(A)选择细分市场,以便进一步
研究,(B)过滤掉不具信息性的片段
数据库搜索,以及c)发现和分析
哪种成分的偏向以周期性间隔出现
而不是连续的残基。自动化的新方法
低复杂度序列的分类及其邻域
已经被开发出来了。
B.丰度和生物特性:约25%
蛋白质数据库中的残基在组成上
有偏差的部分(包括一些已知的长而非球状的部分
区域),大约55%的蛋白质含有一个或
更多这样的细分市场。交错低复杂度序列
在许多领域特别丰富。点缀着
对于许多人来说,低复杂度的序列尤其丰富
真核细胞在形态发生和胚胎发育中的关键蛋白
发育,RNA加工,转录调控,
信号转导与细胞和细胞因子
细胞外结构的完整性。有限的结构
可用于以下低复杂性区域的信息
蛋白质表明它们通常是非球形的,
多态的或移动的。
该项目突出了高丰度和
低复杂性蛋白质片段的生物学重要性。
关于它们的分子结构和动力学的知识
开始出现在一些案例中,但这些都是
少数族裔。这是未来研究的优先领域。
最近发展起来的分析核苷酸的方法
序列揭示了许多新的和错综复杂的
构成特征。这些方法在以下方面很有价值
消除序列数据库搜索中的许多伪像
和比对分析。
英文摘要
The goal of this project is to define, classify and
analyze, using computational analysis, segments of protein
and nucleotide sequences showing compositional bias or
improbably low compositional complexity. In protein
sequences, these include the abundant residue clusters of
predominantly one or a few amino acid types, which commonly
contain homopolymeric tracts or mosaics of these, aperiodic
patterns and sections of low-period repeats. Other common
examples include long non-glomerular domains. The
abundance of biased segments in both amino acid and
nucleotide sequence databases has been determined, and
their properties are being related to evidence of
biological functions.
Different formal definitions of local compositional
complexity were used to make unbiased identification of
low-complexity segments, at different levels of stringency.
Algorithms were refined to (a) select segments for further
study, (b) filter out non-informative segments prior to
database searches, and c) discover and analyze regions in
which compositional bias is present in periodically-spaced
rather than contiguous residues. New methods for automated
classification and neighboring of low-complexity sequences
have been developed.
B. Abundance and biological properties: Approximately 25%
of the residues in protein databases are in compositionally
biased segments (including some known long non-globular
regions) and approximately 55% of proteins contain one or
more such segments. Interspersed low-complexity sequences
are particularly abundant in many segments. Interspersed
low-complexity sequences are particularly abundant to many
eukaryotic proteins crucial in morphogenesis and embryonic
development, RNA processing, transcriptional regulation,
signal transduction and aspects of cellular and
extracellular structural integrity. The limited structural
information available for low-complexity regions of
proteins indicates that they are generally non-globular and
polymorphic or mobile.
The project is highlighting the high abundance and
biological importance of low-complexity protein segments.
Knowledge of their molecular structure and dynamics is
beginning to emerge in a few cases, but these are a
minority. This is a priority area for future research.
The methods recently developed to analyze nucleotide
sequences are revealing many new and intricate
compositional features. These methods are valuable in
eliminating many artifacts in sequence database searches
and alignment analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MOLECULAR NOVELTY IN SEQUENCES OF BACTERIA AND MODEL ORGANISMS
-
批准号:6162793
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:J C WOOTTON
-
依托单位:
MOLECULAR NOVELTY IN SEQUENCES OF BACTERIA AND MODEL ORGANISMS
-
批准号:2578625
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:J C WOOTTON
-
依托单位:
COMPUTER ANALYSIS OF LOW-COMPLEXITY AMINO ACID AND NUCLEOTIDE SEQUENCES
-
批准号:2578624
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:J C WOOTTON
-
依托单位: