课题基金 / 基金详情

ANNOTATING HUMAN GENOME BY MASS SPECTROM

ANNOTATING HUMAN GENOME BY MASS SPECTROM
通过质谱注释人类基因组
批准号:
7355049
负责人:
MARKUS KALKUM
金额:
$0.12万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-03-01 至 2007-02-28

项目摘要

项目成果

MARKUS KALKUM的其他基金

相似基金

相关文献

中文摘要
翻译
本子项目是利用由NIH/NCRR资助的中心赠款提供的资源的众多研究子项目之一。子项目和研究者(PI)可能已经从另一个NIH来源获得了主要资金,因此可以在其他CRISP条目中表示。列出的机构是中心的,不一定是研究者的机构。尽管人类基因组的工作草图已经包含了超过25%的完成序列,但具有正确外显子和外显子/内含子边界的基因的分配仍然是一个主要挑战。如果没有或只有部分c-DNA序列存在,基因预测可以在计算机上进行,例如,使用基因扫描和/或同源性搜索。然而,对不存在的外显子的过度预测是这些算法的一个特点,特别是如果将解析标准设置为检测次优外显子,以便不太可能遗漏真正的外显子(http://genes.mit.edu/Suboptimal.html)。我们开发了一种新的外显子定位策略(使用从蛋白质基因产物生成的MS2数据),可以安全地揭示表达蛋白中存在哪些外显子以及有效的外显子/内含子边界。我们使用ESI-和maldi -离子阱质谱仪生成来自人类蛋白质的蛋白水解肽的高质量MS/MS数据,这些MS/MS数据使用新开发的搜索算法“Sonar”与可用的人类基因组序列相关联。对一个肽的单个可靠命中提供了鉴定序列位于外显子内的证据。使用基因组序列中围绕命中位置的适当区域来生成粗略的基因预测。预测的外显子被交替组装,不同的组装用获得的所有MS/MS数据对同一蛋白质进行搜索。不击中基因组序列但桥接预测外显子的肽导致外显子边界的明确基因注释。作为这种分析的一个例子,我们使用了从人类STAGA复合体中获得的130kDa波段的单次LC-MS/MS运行的数据。检索整个公开可用的人类基因组数据库(截至2001年1月27日),我们确定了9个不同的外显子,3个不同的外显子连接属于TAF2C1基因(TATA box binding protein (TBP)-associated factor, RNA polymerase II, C1)。这些数据跨越了大约100个基因组序列的区域。我们特别获得了7个外显子内肽和3个外显子桥接肽。应该提到的是,TAF2C1的第一个和最后一个外显子(15个外显子)是假定的,并且被基因组序列中的大间隙中断。我们将提供一系列其他示例,并讨论该策略在注释人类基因组序列以及其他应用(如蛋白质的可选剪接变体的定义)中的价值。“Sonar”提供了一种有效的新评分算法,以及一种呈现有意义的MS/MS数据的新方法,使我们能够以高速(600 msec/谱)筛选基因组数据库中的蛋白质。我们得出的结论是,以这种方式使用的质谱法对于注释人类基因组具有重要的实用性。一篇描述这项工作的论文正在准备中。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. Although the working draft of the human genome already contains more than 25% finished sequence, the assignment of genes with their correct exons and exon/intron-boundaries remains as a major challenge. If none or just partial c-DNA sequences are present, gene-prediction can be performed in-silico, e.g., using genscan and/or homology searches. However, overprediction of nonexistent exons is a feature of these algorithms, especially if the parsing criteria are set to detect suboptimal exons so that real exons are less likely to be missed (http://genes.mit.edu/Suboptimal.html). We developed a new exon mapping strategy (using MS2 data generated from the protein gene products) that safely reveals which exons are present in an expressed protein as well as valid exon/intron boundaries. We generate high quality MS/MS data of proteolytic peptides derived from human proteins using ESI- and MALDI-ion trap MS. This MS/MS data is correlated with the available human genome sequence using a newly developed search algorithm called "Sonar". A single confident hit for a peptide provides evidence that the identified sequence lies within an exon. An appropriate region of the genome sequence surrounding the hit position is used to generate a rough gene prediction. The predicted exons are alternatively assembled and the different assemblies are searched with all MS/MS data obtained for the same protein. Peptides that do not hit the genomic sequence but bridge predicted exons lead to unambiguous gene annotation of exon boundaries. As an example of such an analysis, we used data from a single LC-MS/MS run of a 130kDa band obtained from the human STAGA complex. Searching the entire publicly available human genome database (as of January 27, 2001) we identified 9 different exons with 3 distinct exon junctions that belong to the TAF2C1 gene (TATA box binding protein (TBP)-associated factor, RNA polymerase II, C1). The data spanned a region of ~100 kbases of genomic sequence. In particular we obtained 7 peptides within exons and 3 exon-bridging peptides. It should be mentioned that the first and the last exon of TAF2C1 (15 exons) are putative and are interrupted by large gaps in the genomic sequence. We will provide a series of other examples and discuss the value of this strategy for annotating the human genome sequence, as well as for other applications such as the definition of alternatively spliced variants of proteins. "Sonar" provides an effective new scoring algorithm plus a novel way of presenting meaningful MS/MS data, enabling us to screen proteins against genomic databases at high speed ( 600 msec/spectrum). We conclude that mass spectrometry used in this way can be of significant utility for annotating the human genome. A paper describing this work is in preparation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Diagnostic assay systems for Botulinum Neurotoxins
Diagnostic assay systems for Botulinum Neurotoxins
Fluorescent and bioluminescent assays for botulinum neurotoxin
  • 批准号:
    8260257
  • 项目类别:
  • 资助金额:
    $40.35万
  • 财政年份:
    2011
  • 负责人:
    MARKUS KALKUM
  • 依托单位:
Diagnostic assay systems for Botulinum Neurotoxins
国内基金
海外基金
靶向Human ZAG蛋白的降糖小分子化合物筛选以及疗效观察
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    胡文静
  • 依托单位:
HBV S-Human ESPL1融合基因在慢性乙型肝炎发病进程中的分子机制研究
  • 批准号:
    81960115
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    34.0万元
  • 批准年份:
    2019
  • 负责人:
    江建宁
  • 依托单位:
基于自适应表面肌电模型的下肢康复机器人“Human-in-Loop”控制研究
  • 批准号:
    61005070
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2010
  • 负责人:
    李庆玲
  • 依托单位: