课题基金 / 基金详情

FINDING PROTEIN SEQUENCE MOTIFS--METHODS AND APPLICATIONS

FINDING PROTEIN SEQUENCE MOTIFS--METHODS AND APPLICATIONS
寻找蛋白质序列基序——方法和应用
批准号:
6162801
负责人:
E V KOONIN
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

E V KOONIN的其他基金

相似基金

相关文献

中文摘要
翻译
随着序列信息的快速增长,大大增加了 取代实验数据的累积速度 蛋白质功能,敏感方法的作用 蛋白质序列分析,包括检测 微妙但功能上重要的主题,一直是 越来越多。该项目的目标包括 一种描述蛋白质的连贯策略的发展 最终,超家族和预测蛋白质功能 旨在建立一个全面的数据库 蛋白质功能基序。使用的方法包括 使用单个序列进行序列数据库搜索( BLAST和FASTA家庭的计划)和多个 序列比对(构建的HMMer程序包 多重比对的隐马尔可夫模型及其应用 它们用于数据库筛选);检测方法 蛋白质序列中的基序,包括在 该项目的早期阶段(过去的计划、CAP、MOST、 吉布斯);多序列比对方法(程序 金刚鹦鹉;蛋白质的分离方法 预测的球状和非球状结构域的序列 (具有可变参数的程序SEG).方法 蛋白质二级结构的预测(程序PHD, 线圈)、跨膜结构域(PHDhtm)和信号肽 (Signalp);预测DNA编码区的方法 基于非齐次马尔可夫模型(GeneMark);方法 用于根据序列相似性(CLU)对蛋白质进行聚类。 这些方法被组合成一种序列分析策略 设计主要是为了有效地分析 大的、多结构域的蛋白质序列,组成 大多数与人类有牵连的基因产物 疾病。首先将蛋白质序列划分为 假定的球状和非球状区域,之后 数据库搜索是分别与 单个球状域的序列使用 传递性BLAST搜索和Motif的组合 分析。除了通用序列之外 数据库,独立的、较小的数据库被构建 利用蛋白质功能和/或系统发育的信息 起源。两个大型数据集,即基因的产物 参与动物开发和产品的 对定位克隆的人类疾病基因进行了分析 使用这些方法。各种以前的 未确定特征,但具有潜在的重要功能 发现了结构域和基序。两个重要的例子 包括一个假定的FAD结合结构域 二核苷酸结合修饰的脉络膜血症蛋白 共识阻止了它之前的检测,以及一个 指定为BRCT的域,该域在多个 参与DNA损伤反应细胞周期的蛋白质 检查点,包括人类BRCA1基因的产物 与遗传性乳腺癌和卵巢癌有关。
英文摘要
With the rapid growth of sequence information which greatly supersedes the rate of accumulation of experimental data on protein functions, the role of sensitive methods for protein sequence analysis, including the detection of subtle but functionally important motifs, is constantly increasing. The goals of this project include the development of a coherent strategy for delineating protein superfamilies and predicting protein function, eventually aiming at the construction of a comprehensive database of protein functional motifs. The methods used included sequence database search with individual sequences (the programs of the BLAST and FASTA families) and multiple sequence alignments (HMMer program package that builds Hidden Markov Models from multiple alignments and applies them for database screening); methods for detection of motifs in protein sequences, including those developed at an earlier stage of this project (programs PAST, CAP, MoST, GIBBS); multiple sequence alignment methods (programs MACAW, CLUSTALW); methods for partitioning protein sequences into predicted globular and non-globular domains (program SEG with varying parameters); methods for prediction of protein secondary structure (programs PHD, COILS), transmembrane domains (PHDhtm), and signal peptides (Signalp); a method for prediction of coding regions in DNA based on non-homogeneous Markov models (GeneMark); methods for clustering proteins by sequence similarity (CLUS). These methods were combined in a sequence analysis strategy designed primarily in order to efficiently analyze the sequences of large, multidomain proteins which comprise the majority of the products of genes implicated in human diseases. The protein sequences were first partitioned into putative globular and non-globular domains, after which database searches were conducted separately with the sequences of individual globular domains using a combination of transitive BLAST searches and motif analysis. In addition to general purpose sequence databases, separate, smaller databases were constructed using information on protein function and/or phylogenetic origin. Two large data sets, namely the products of genes involved in animal development and the products of positionally cloned human disease genes, were analyzed using these approaches. A variety of previously uncharacterized but potentially functionally important domains and motifs were discovered. Two important examples include a putative FAD-binding domain in the human choroideremia protein with a modified dinucleotide-binding consensus which prevented its previous detection,and a domain designated BRCT, which is conserved in a number of proteins involved in DNA damage-responsive cell cycle checkpoints, including the product of the human BRCA1 gene implicated in hereditary breast and ovarian cancers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
COMPUTER-ASSISTED DISSECTION OF ROLLING CIRCLE DNA REPLICATION
  • 批准号:
    3845128
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    E V KOONIN
  • 依托单位:
GENOME ORGANIZATION AND EVOLUTION OF RNA VIRUSES
  • 批准号:
    3845123
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    E V KOONIN
  • 依托单位:
COMPUTER-ASSISTED STUDY OF FUNCTIONS AND EVOLUTION OF LARGE DNA VIRUS GENOMES
  • 批准号:
    3845124
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    E V KOONIN
  • 依托单位:
COMPREHENSIVE COMPUTER ANALYSIS OF E COLI GENES
  • 批准号:
    3781286
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    E V KOONIN
  • 依托单位:
海外基金