课题基金 / 基金详情

项目摘要

项目成果

David B Jaffe的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):读取生物体基因组DNA序列的能力已经彻底改变了生物学和生物医学研究。准确的序列数据集合是至关重要的,因为它们为所有后续工作提供了基础。利用基于毛细管的测序技术,生成了许多基因组的高质量草稿。在过去的几年里,大规模平行测序技术已经将测序成本降低了1000倍,但这些技术的测序结果比毛细管测序结果更短,更不准确,因此更难组装,特别是对于大型基因组。我们最近展示了大量并行数据的集合,这些数据开始接近毛细管数据的质量。这些基因组的组装是非常高质量的(“完成的”)组装已经可用,因此我们不仅能够严格评估组装的质量,还能够系统地诊断它们的缺陷。此外,我们观察到,在几乎所有情况下,有缺陷的基因座都有足够的覆盖范围,只要有合适的算法,原则上就可以正确地组装它们。在此基础上,我们提出了一项研究计划,以开发用于创建前所未有质量的组装的计算方法:在我们的第一个目标中,我们建议开发方法来实现高质量的新基因组草图组装。我们的目标是达到并超过毛细管测序所达到的质量水平。在我们的第二个目标中,我们将开发方法来实现超高质量的人类基因组组装。为此,我们将利用现有的人类参考序列和其他个体的参考序列,包括我们将创建的参考序列。通过这种方式,我们的目标是为参考序列中表示的区域实现接近完成的质量(主要是通过“重测序”方法),同时(通过从头开始的方法)捕获参考序列中不存在的区域。因此,我们的目标是尽可能地为每个人的基因组提供最好的代表。我们注意到,随着成本的下降,这很可能成为患者的“标准护理”。在我们的第三个目标中,我们超越现有的数据,着眼于下一代测序技术,使用非常长的和“频闪”读取来组装非常困难的区域。这些困难区域包括片段重复,这是进化热点,与许多疾病有关,除了使用非常昂贵的逐个克隆测序的方法外,目前的方法无法获得。最后,我们的第四个目标是使社区可以使用汇编方法。在这里,我们的目标是使一系列用户(包括个人研究者)尽可能容易地匹配基因组组装专家可以实现的结果。简而言之,通过我们的四个目标,我们将使社区能够使用最低成本的数据实现最高的装配质量。因此,我们预计我们的工作将推动对生物学和人类疾病具有重要意义的广泛研究。
英文摘要
DESCRIPTION (provided by applicant): The ability to read the DNA sequence of an organism's genome has revolutionized biology and biomedical research. Accurate assemblies of sequence data are critical because they provide the foundation for all subsequent work. Using capillary-based sequencing technology, high quality drafts were generated for many genomes. Over the past several years, massively parallel sequencing technologies have lowered sequencing cost by 1000-fold, but the reads from these technologies are shorter and less accurate than the capillary reads, hence harder to assemble, particularly for large genomes. We have recently demonstrated assemblies of massively parallel data that begin to approach the quality of those from capillary data. These assemblies were of genomes for which exceptionally high-quality ('finished') assemblies were already available, and we were thus able not only to rigorously assess the quality of our assemblies, but also to systematically diagnose their defects. Moreover we observe that in almost all cases, defective loci have enough coverage that they could in principle be assembled correctly, provided that the right algorithms were available. On this basis we have proposed a research program to develop computational methods for the creation of assemblies of unprecedented quality: In our first aim we propose to develop methods to achieve high quality draft assemblies of new genomes. Here our objective is to reach and exceed the level of quality that had been achieved using capillary sequencing. In our second aim we will develop methods to achieve ultra high quality assemblies of human genomes. To do this we will leverage the existing human reference sequence and reference sequences of other individuals, including those that we would create. In this way we aim to achieve near-finished quality for regions represented in the reference sequences (essentially via 'resequencing' methods), and at the same time (by de novo methods) capture those regions that are not present in the reference sequences. Our aim is thus to produce the best possible representation of each individual's genome. We note that as costs drop, this is likely to become 'standard of care' for patients. In our third aim, we look beyond existing data, to the next generation of sequencing technologies, to assemble very hard regions using very long and 'strobe' reads. These hard regions include segmental duplications, which are evolutionary hotspots, associated with many diseases, and inaccessible to current methods, except those using very expensive clone-by-clone sequencing. Finally our fourth aim is to make assembly methods accessible to the community. Here our goal is to make it as easy as possible for a range of users (including individual investigators) to match the results achievable by genome assembly experts. In short, through our four aims, we will enable the community to achieve the highest possible assembly quality using the lowest cost data. We thus anticipate that our work will advance a broad range of investigations of importance to biology and human disease. PUBLIC HEALTH RELEVANCE: This grant will develop better methods for completely and accurately determining the genome sequence of an organism, in particular producing precise representations of complex and repetitive regions of genomes. This will advance a broad range of investigations in genome evolution, cancer and human genetic disease.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Whole-genome shotgun sequencing strategy and assembly
  • 批准号:
    7892886
  • 项目类别:
  • 资助金额:
    $5.92万
  • 财政年份:
    2009
  • 负责人:
    David B Jaffe
  • 依托单位:
Whole-genome shotgun sequencing strategy and assembly
  • 批准号:
    7917781
  • 项目类别:
  • 资助金额:
    $77.23万
  • 财政年份:
    2009
  • 负责人:
    David B Jaffe
  • 依托单位:
Whole-genome shotgun sequencing strategy and assembly
  • 批准号:
    8334610
  • 项目类别:
  • 资助金额:
    $76.0万
  • 财政年份:
    2005
  • 负责人:
    David B Jaffe
  • 依托单位:
Whole-genome shotgun sequencing strategy and assembly
海外基金