课题基金 / 基金详情

项目摘要

项目成果

David E. Hill的其他基金

相似基金

相关文献

中文摘要
翻译
摘要 阐明基因组的编码潜力得益于准确的基因组序列和广泛的 转录组测序,以实现蛋白质编码序列(CDS)或开放阅读框架的详细模型 (ORF)。尽管至少为每个蛋白质编码分配了一个可靠的全长转录本模型 基因,由于i)表达水平的巨大差异,大多数可供选择的亚型仍未确定 在共同基因表达的异构体之间,以及ii)获得全长(FL)转录本的困难 序列。此外,两年的成绩单总数仍有很大差异。 注解数据库和有实验证据的注解FL成绩单的编号。 编码文本的光谱由一个巨大但有限的多维“同形-空间”组成:i) 基因,ii)组织和细胞类型,iii)发育和时间,iv)疾病,v)对刺激的反应。就像 不同细胞和组织的表达水平不同,选择性剪接转录本的相对丰度也不同。 如果没有经验知识,对人类基因组的全面、功能的理解是不可能的 对整个编码功能蛋白质的完整注释。 从历史上看,基因注释主要由INSDC数据库中的EST和mRNAs支持,而 自动化的注释方法正在应用于整个基因组和转录本。但是,当前 自动批注不能提供与手动批注相同质量的数据。敏感性和 专用性降低,捕获的功能批注较少,并且所有自动化方法都缺乏 手动注释器引入额外的正交数据类型和对科学文献的解释,但是 手动标注是高度劳动密集型的。GENCODE版本v36代表了对近10个 百万个EST、cDNA和蛋白质同源序列。在给定预期数据量的情况下,使用单个实验 产生的数据比整个INSDC目录都多,目前的手动注释方法不能扩展。这个 长转录组测序方法的出现提供了历史数据类型的替代 基因和转录本注释的好处。然而,海量的数据量已经在 存放在公共数据档案中的数据超过了人工管理能力,需要实施自动化 在不影响批注质量的情况下提供解决方案。此外,由于非靶向测序方法非常 在他们发现不太丰富的转录本方面效率低下,产生的大多数序列数据给了我们非常 对可发现的记录多样性缺乏洞察力。为了克服这些挑战,我们两个小组分别 联合起来增加经过充分实验验证的全长人类蛋白质编码转录本的目录。 这项建议侧重于实验方法的整合,这将提供一个全面的 人类蛋白质编码转录本的列举--随着人类蛋白质编码转录组的发展而产生的“参考人类转录组” 自动注释管道,允许将此资源整合到GENCODE基因注释中。
英文摘要
Abstract Elucidating the coding potential of the genome has benefited from accurate genome sequences and extensive transcriptome sequencing to allow detailed models for protein-coding sequences (CDSs) or open reading frames (ORFs). Although at least one reliable full-length transcript model has been assigned for every protein-coding gene, the majority of alternative isoforms remains uncharacterized due to i) vast differences of expression levels between isoforms expressed from common genes, and ii) the difficulty of obtaining full-length (FL) transcript sequences. Furthermore, there remains a large discrepancy between the total number of transcripts in annotation databases and the number for which there is an annotated FL transcript with experimental evidence. The spectrum of encoded transcripts comprises a vast but finite “isoform-space” with multiple dimensions: i) genes, ii) tissues and cell types, iii) development and time iv) disease, and v) response to stimuli. Just as expression levels vary across cells and tissues, so can the relative abundance of alternatively spliced transcripts. Full, functional understanding of the human genome will not be possible without empirical knowledge and complete annotation of the entire complement of encoded functional proteins. Historically, gene annotation was supported predominantly by ESTs and mRNAs from INSDC databases while automated approaches to annotation are being applied to whole genomes and transcriptomes. However, current automated annotation does not provide the same quality data as does manual annotation. Sensitivity and specificity are reduced, less functional annotation is captured, and all automated methods lack the capacity of a manual annotator to introduce additional orthogonal data types and interpretation of the scientific literature, but manual annotation is highly labor-intensive. GENCODE release v36 represents the interpretation of nearly 10 million EST, cDNA and protein homologies. Given the anticipated volumes of data, with single experiments producing more data than the entire INSDC catalogue, current methods of manual annotation do not scale. The emergence of long transcriptomic sequencing methods provides for the replacement of historical data types to the benefit of gene and transcript annotation. However, the massively greater data volumes already being deposited in public data archives exceed manual curation capability, demanding implementation of automated solutions without compromising annotation quality. Furthermore, as untargeted sequencing approaches are very inefficient in their discovery of less abundant transcripts, the majority of sequence data generated gives us very little insight into discoverable transcript diversity. To overcome these challenges, our two respective groups have joined forces to increase the catalog of fully experimentally verified full length human protein-coding transcripts. This proposal focuses on the integration of experimental approaches that will provide a comprehensive enumeration of human protein-coding transcripts, a “Reference Human Transcriptome” with the development of an automated annotation pipeline to allow the integration of this resource into GENCODE gene annotation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Generating a full-length reference transcriptome for human protein-coding genes
  • 批准号:
    10687972
  • 项目类别:
  • 资助金额:
    $66.35万
  • 财政年份:
    2022
  • 负责人:
    David E. Hill
  • 依托单位:
The 6th ORFeome Meeting: ORFeomes and Systems
  • 批准号:
    7225045
  • 项目类别:
  • 资助金额:
    $0.8万
  • 财政年份:
    2006
  • 负责人:
    David E. Hill
  • 依托单位:
Mapping the first half of the REFERENCE human binary protein interactome
  • 批准号:
    8518435
  • 项目类别:
  • 资助金额:
    $176.78万
  • 财政年份:
    1998
  • 负责人:
    David E. Hill
  • 依托单位:
Mapping the Human Binary Interactome Network
  • 批准号:
    7688648
  • 项目类别:
  • 资助金额:
    $184.19万
  • 财政年份:
    1998
  • 负责人:
    David E. Hill
  • 依托单位:
海外基金