课题基金 / 基金详情

项目摘要

项目成果

Paul Flicek的其他基金

相似基金

相关文献

中文摘要
翻译
资源信息-项目摘要 GENCODE资源的创建、发展和维护需要 遵守和优化规定的流程,确保创建基因组注释 现在和将来的标准将永远与过去相同或更好。 已经被创建。GENCODE资源也必须适应新技术 以及随着基因组学领域的发展而出现的机会。的主要目的 GENCODE资源用于确保注释的质量控制(QC)和数据验证。 Ensembl将比较GENCODE基因集与其他基因集(例如UniProt),以检查 缺失基因或转录本; CNIO将验证编码基因; CNIO/CNIC蛋白质组学 pipeline将验证基因模型; CNIO/CNIC将对 蛋白质组学数据。项目的稳定性将通过一个维护良好的计算 基础设施,充分的质量控制流程,以确保最高的质量,以及 定期发布高价值格式的免费注释。的注释策展 人类和小鼠的部分转录本模型,特别是现有的人类部分转录本模型 将扩展到全长,扩展人类lncRNA注释,以及 完成鼠标注释的初始完整过程。GENCODE将包含个人 基因组表示和由可用的人类变异数据表示的群体数据, 序列水平(例如1000个基因组)和转录组水平(例如GTEx),以及 由WTSI领导的小鼠基因组计划产生的16个小鼠品系基因组。 将对来自个人和群体的数据进行注释。个人基因组资源将是 这将产生一个准确的代表性的个人的基因集。两个试点 这些项目将有助于确定支持未来GENCODE注释的最有效方式。的 第一个试点项目将利用GENCODE在开发人口参考基因组方面的经验 图表,以引导基于人群的基因组的可扩展和潜在的通用方法 注释。第二个试点项目将侧重于将监管区域与受监管区域连接起来, 基因. GENCODE将增强当前对具有调控元件的基因的注释 使得注释取决于组织和细胞类型。手动注释的需求 跨菌株和物种的转录本的数量可能超过GENCODE提供这种能力的能力, 服务,因此,一个系统,使提交注释 将开发数据。所述措施将确保2020年的GENCODE将 对基因组学的研究和临床应用来说比今天更有价值。
英文摘要
RESOURCE INFORMATICS – PROJECT SUMMARY The creation, advancement and maintenance of the GENCODE resource requires both adherence to and optimization of defined processes that ensure the genome annotation created now and in the future will always be of the same or better standard compared to what has already been created. The GENCODE resource must also be attuned to the new technologies and opportunities that arise as the field of genomics evolves. A primary objective of the GENCODE resource is to ensure quality control (QC) and data validation of annotations. Ensembl will compare the GENCODE gene set to other gene sets (e.g. UniProt) to check for missing genes or transcripts; CNIO will validate the coding genes; the CNIO/CNIC proteomics pipeline will validates the gene models; CNIO/CNIC will perform manual verification for QC of proteomics data. Project stability will be ensured through a well-maintained computational infrastructure, adequate QC processes that will ensure the highest possible quality, as well as regular releases of freely available annotation in high value formats. The annotation curation for human and mouse will be completed, in particular the existing human partial transcript models will be extended to full length, expanding the human lncRNA annotation, as well as the completion of the initial full pass of the mouse annotation. GENCODE will incorporate individual genome representation and population data represented by available human variation data at both the sequence level (e.g. 1000 Genomes) and at the transcriptomic level (e.g. GTEx), and by the 16 mouse strain genomes produced by the Mouse Genomes Project led by the WTSI. Data from individuals and populations will be annotated. A personal genome resource will be developed, which will produce an accurate representation of an individual's gene set. Two pilot projects will help to define the most effective way to support future GENCODE annotations. The first pilot project will use GENCODE's experience in developing population reference genome graphs to pilot a scalable and potentially universal approach to population based genome annotation. The second pilot project will focus on connecting regulatory regions to regulated genes. GENCODE will enhance the current annotation of genes with their regulatory elements so that the annotation is dependent on tissue and cell type. The demand for manual annotation of transcripts across strains and species may outstrip GENCODE's ability to provide such services via existing mechanisms, therefore a system to enable the submission of annotated data will be developed. The described measures will ensure that GENCODE in 2020 will be significantly more valuable for research and clinical applications in genomics than today.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
The WashU-UCSC-EBI Human Genome Reference Center."
  • 批准号:
    10419218
  • 项目类别:
  • 资助金额:
    $25.0万
  • 财政年份:
    2021
  • 负责人:
    Paul Flicek
  • 依托单位:
Enabling Comparative Pangenomics
Enabling Comparative Pangenomics
Enabling Comparative Pangenomics
海外基金