课题基金 / 基金详情

项目摘要

项目成果

LINCOLN D. STEIN的其他基金

相似基金

相关文献

中文摘要
翻译
描述:modENCODE项目是果蝇和蠕虫基因组测序的关键后续,将对我们对包括人类在内的所有高等真核生物生物过程的理解产生巨大影响。为了管理modENCODE将产生的各种大规模数据集,我们建议创建一个数据协调中心(DCC)来跟踪数据,将其与其他信息来源整合,并使其及时开放地提供给研究社区。该提案汇集了四个具有高度相关背景的小组:Micklem小组通过其在InterMine系统和FlyMine数据库上的工作,在将不同类型的数据集成到高性能数据挖掘系统方面拥有丰富的经验。Stein和Lewis团队为该项目带来了对线虫和黑胃线虫基因组、它们的试剂和研究群体的熟悉,并通过他们与WormBase和FlyBase数据库的合作,很好地定位了与这些mod的联系。肯特集团负责人类ENCODE试点项目的DCC,并在开发和管理此类项目方面拥有广泛的实践知识。我们将在CSHL和伯克利组建一个由三名数据管理人员组成的团队,他们具有秀丽隐杆线虫和/或D. melanogaster的生物信息学背景。管理人员将与他们在数据提供商站点的联系人联系,以确定数据文件格式、里程碑和数据集的质量控制程序。他们还将与NCBI的代表联络,协调modENCODE活动与GenBank和GEO的主要数据存储库。数据提供者将把他们的数据集上传到一个暂存服务器上,这样他们就可以在GBrowse基因组浏览器的一个实例上预览他们的数据。数据管理员将在批准将数据转移到生产数据库之前对其进行质量控制。数据将使用InterMine集成到生产数据库中,并每月向公众发布一次。研究人员将能够通过GBrowse基因组浏览器、批量下载以及通过InterMine和BioMart数据仓库系统介导的复杂查询和报告访问数据。提议的DCC使用的所有主要软件系统将基于来自通用模型生物数据库(GMOD)、人类ENCODE和其他来源的开源工具。在整个项目中,Lewis和Stein将与FlyBase和/或WormBase密切合作,以确保modENCODE收集的数据成为相关模式生物数据库的组成部分。此外,在项目的最后一年,我们将投入数据管理器的大部分工作,将数据从modENCODE转移到mod中。
英文摘要
DESCRIPTION: The modENCODE project is a key sequel to the sequencing of the fly and worm genomes, and will have an enormous impact on our understanding of biological processes in all higher eukaryotes, including human. In order to manage the diverse, large-scale datasets that will be produced by modENCODE, we propose to create a data coordinating center (DCC) to track the data, integrate it with other information sources, and make it available to the research community in a timely and open fashion. This proposal brings together four groups with highly relevant backgrounds: The Micklem group, through its work on the InterMine system and FlyMine database, has extensive experience in integrating diverse types of data into high-performance data mining systems. The Stein and Lewis groups bring to the project an intimate familiarity with the C. elegans and D. melanogaster genomes, their reagents and research communities, and are well-positioned by their work with the WormBase and FlyBase databases to liaise with those MODs. The Kent group is responsible for the DCC for the Human ENCODE pilot project, and has extensive practical knowledge of developing and managing projects of this sort. We will assemble a team of three data managers stationed at CSHL and at Berkeley, who have a background in the bioinformatics of C. elegans and/or D. melanogaster. The managers will liaise with their contacts at the data provider sites to determine data file formats, milestones and quality control procedures for their datasets. They will also liaise with representatives from NCBI to coordinate modENCODE activities with the primary data repositories at GenBank and GEO. Data providers will upload their data sets to a staging server where they will be able to preview their data on an instance of the GBrowse genome browser. The data managers will QC the data before approving its transfer to the production database. Data will be integrated in the production database using InterMine, and from there released to the public on a monthly schedule. Researchers will be able to access the data via the GBrowse genome browser, bulk downloads, and via complex queries and reports mediated by InterMine and the BioMart data warehousing system. All major software systems used by the proposed DCC will be based on open source tools from the Generic Model Organism Database (GMOD), human ENCODE, and other sources. Throughout the project, Lewis and Stein will work close with FlyBase and/or WormBase to ensure that data collected by modENCODE becomes an integral part of the relevant model organism database. In addition we will dedicate a significant part of a data manager's effort to transfer data from modENCODE into the MODs during the last year of the project.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Overture: A Multi-Scale Data Platform for Cancer Genomics Research
A Data Coordinating Center for modENCODE
A Data Coordinating Center for modENCODE
Reactome: An Open Knowledgebase of Human Pathways
海外基金