课题基金 / 基金详情

项目摘要

项目成果

TIMOTHY J DURFEE的其他基金

相似基金

相关文献

中文摘要
翻译
我们正在进入个人基因组学的新时代,在这个时代,个人的基因组序列将被用来 识别疾病易感性,改进诊断和更好地治疗疾病,以及跨 队列和群体,以确定新的生物标志物和任何表型背后的因果突变。尽管 将短读下一代测序(NGS)数据映射到参考数据上的巨大成功 基因组(重测序)在识别新基因组中的遗传变异时,先天缺乏长距离 连接性和参考诱导的偏差使获得完整的单倍型阶段性基因组 非常困难。新兴的长读技术正开始通过以下方式解决这一关键缺陷 个体基因组的直接从头组装。然而,最初的从头装配通常包括 数以千计的无序重叠群,需要大量的组装后处理才能产生成品 可以有效地挖掘基因内容和变异的序列。因此,迫切需要 集成的、可扩展的装配后软件,可1)自动组织、加入初始重叠群并对其进行阶段化 到完整的单倍型序列,2)支持可选的NGS和/或手动抛光,以及3)提供初始 这些序列的自动注释。目前,这样的软件还不存在,相反,用户必须 拼凑了一系列令人困惑的难以使用的、特定于任务的开放源码程序。 DNAStar的装配后编辑程序SeqMan Pro(SMP)在整理细菌方面具有成熟的历史 大小的基因组,尽管它目前缺乏可扩展性和处理人类所需的所有功能 基因组大小的问题。此Fast Track提案的主要目标是创建一个完全可扩展的版本 用于从头组装的大的真核基因组的自动完成和注释的SMP,同时还 在需要时提供手动编辑平台。在第一阶段,我们将开发两个关键原型:1)a 新的部件文件格式eBAM,它可以与BAM格式相互转换,但也可以像我们的 Sqd文件和2)一个快速参考辅助的重叠群搭建工具,改编自我们专有的磁盘排序 比对(DSA)算法。在此基础上,我们将通过以下方式完成SMP二期改造:1) 优化eBAM格式以获得最佳编辑性能,2)构建新的位版本的SMP编辑 该引擎集成了大型真核生物组装后完成所需的附加功能 基因组包括基于DSA的自动化支架和阶段感知缺口填充、重叠群连接和 单倍型精化,3)创建新的基于DSA的基因组比对器,用于快速比对已完成的序列 一个带注释的参考基因组,再加上一个新的特征转移和分析模块,将允许 完成基因组的初始注释,以及变异体的编目及其对原生和 参考坐标。参考坐标的包含使得新基因组中的变异很容易 与众多在线知识库资源提供的丰富信息相关联。
英文摘要
We are entering a new era of personal genomics where an individual's genome sequence will be used to identify disease susceptibility, improve diagnosis and better treat illnesses as well as be combined across cohorts and populations to identify new biomarkers and causal mutations underlying any phenotype. Despite the tremendous success of mapping short read next-generation sequencing (NGS) data onto a reference genome (resequencing) in identifying genetic variation in a new genome, the inherent lack of long range connectivity together with reference-induced biases make obtaining complete haplotype-phased genomes exceedingly difficult. Emerging long read technologies are beginning to address this critical shortcoming by direct de novo assembly of an individual's genome. However, initial de novo assemblies typically consist of many thousands of unordered contigs that require extensive post-assembly processing to produce finished sequences that can be effectively mined for genetic content and variation. Thus, there is an urgent need for integrated, scalable post-assembly software that 1) automatically organizes, joins and phases the initial contigs into complete haplotype sequences, 2) supports optional NGS and/or manual polishing and 3) provides initial automated annotation of those sequences. Currently, such software does not exist and instead users must cobble together a confusing array of difficult-to-use, task-specific pieces of open source programs. DNASTAR's post-assembly editing program, SeqMan Pro (SMP), has a proven history in finishing bacterial sized genomes although it currently lacks the scalability and all the needed functionality to tackle human genome sized problems. The primary goal of this Fast Track proposal is to create a fully scalable version of SMP for the automated finishing and annotation of de novo assembled large eukaryotic genomes while also providing a manual editing platform when needed. During Phase I, we will develop two key prototypes: 1) a new assembly file format, eBAM, which is interconvertible with the BAM format, but also is editable like our SQD files and 2) a rapid reference-assisted contig scaffolding tool adapted from our proprietary Disk Sort Alignment (DSA) algorithm. With that foundation, we will complete the transformation of SMP in Phase II by: 1) refining the eBAM format for optimal editing performance, 2) building a new 64-bit version of the SMP editing engine that incorporates the additional functionality necessary for post-assembly finishing of large eukaryotic genomes including automated DSA-based scaffolding and phase-aware gap filling, contig joining and haplotype refinement, 3) creating a new DSA-based genome aligner for rapidly aligning a finished sequence to an annotated reference genome which together with 4) a new feature transfer and analysis module, will permit initial annotation of the finished genome along with a cataloging of variants and their impact in both native and reference coordinates. Inclusion of the reference coordinates allows variants in the new genome to be easily associated with the wealth of information available through the numerous online knowledgebase resources.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Long read based sequencing software for the comprehensive analysis of clinical samples
  • 批准号:
    10009727
  • 项目类别:
  • 资助金额:
    $75.0万
  • 财政年份:
    2020
  • 负责人:
    TIMOTHY J DURFEE
  • 依托单位:
Scalable post-assembly editing software for finishing and annotating personal genomes
  • 批准号:
    9883809
  • 项目类别:
  • 资助金额:
    $75.0万
  • 财政年份:
    2018
  • 负责人:
    TIMOTHY J DURFEE
  • 依托单位:
Complete genome de novo assembly software for the emerging long read sequencing era
  • 批准号:
    9255092
  • 项目类别:
  • 资助金额:
    $74.98万
  • 财政年份:
    2017
  • 负责人:
    TIMOTHY J DURFEE
  • 依托单位:
Complete genome de novo assembly software for the emerging long read sequencing era
  • 批准号:
    9747613
  • 项目类别:
  • 资助金额:
    $6.7万
  • 财政年份:
    2017
  • 负责人:
    TIMOTHY J DURFEE
  • 依托单位:
海外基金