课题基金 / 基金详情

Read-to-contig alignments for de novo genome assembly and annotation

Read-to-contig alignments for de novo genome assembly and annotation
用于从头基因组组装和注释的读取到重叠群比对
批准号:
RGPIN-2014-05112
负责人:
Birol, Inanc
金额:
$2.57万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2014
资助国家:
加拿大
项目状态:
已结题
起止时间:
2014-01-01 至 2015-12-31

项目摘要

项目成果

Birol, Inanc的其他基金

相似基金

相关文献

中文摘要
翻译
这项拟议的研究是关于建立分析DNA的计算技术。DNA是由四种可能的核苷酸(NT)组成的序列:A、C、G和T。在过去的十年里,“读取”DNA序列的技术发生了一场革命,在生命科学的许多领域都有应用。来自高通量测序(HTS)平台的数据达到数亿次“读取”,其中每次读取代表75-300个核苷酸的DNA(人类基因组--我们DNA的总和--大约30亿个核苷酸长)。随着测序技术的发展,解释这些海量的短读数是一个持续的挑战。有两种流行的分析方法来处理HTS读取:基于对齐的方法和基于汇编的方法。第一种使用参考基因组,其DNA序列是从先前对相同或密切相关物种的研究中得知的。在这种方法中,通过搜索读出和参考之间的序列相似性的过程,将读出与参考基因组比对。第二种方法是一种数据驱动的方法,不假定与任何给定的基因组相似。取而代之的是,它重建了由DNA从头开始代表的基因组。这是一种较少偏见的方法,给出了更真实的基因组表示,特别是如果与参考基因组序列相比有重排,或者如果没有参考可用。比罗尔实验室开发了从头组装算法和下游分析工具,并将它们应用于人类健康和其他领域的一些非常引人注目的项目。在拟议的工作中,团队将专注于对齐技术,作为支持这个非常成功的基于组装的分析平台的一种方式。为了与HTS技术发展过程中读取长度和数据量的变化相匹配,已经多次解决了读取对齐问题。然而,高效和准确地将读数与新组装的基因组进行比对是一个尚未得到满足的需求。通用阅读比对算法假设目标序列由少量的长序列组成,本质上是染色体。相比之下,从头开始组装过程的结果通常是数十万件。这给通用校准器带来了问题,我们将通过为这一特定需求开发算法来解决这一问题。我们将特别关注我们算法的可扩展性,以适应不断增长的数据量,我们将通过构建类似于谷歌等互联网搜索引擎所使用的并行处理算法来实现这一点。当一个新物种的基因组被测序和组装时,一个重要的任务是对它的基因进行“注释”--即标记它们在基因组中的位置,以及它们是如何构成的。我们还注意到这一领域的一个重要差距,因为目前的比对技术是为前几代测序平台开发的,已经超过了它们的限制,无法支持来自新测序项目的数据。(这样一个流行的工具,exonerate,仍在大量使用,但它不再由开发人员实验室维护。)我们建议建立一个替代这些工具,并为社区提供持续的支持。随着测序技术的使用进一步渗透到生命科学中,迫切需要高质量的计算工具来及时分析大量数据。所描述的对准技术的发展将提高从头组装及其注释的效率和准确性。
英文摘要
The proposed research is about building computational technologies to analyze DNA. DNA is composed of sequences of four possible nucleotides (nt): A, C, G and T. The last decade witnessed a revolution in technologies that “read” DNA sequences, with applications in many areas of life sciences. Data from high throughput sequencing (HTS) platforms reach hundreds of millions of “reads”, where each read represents 75-300 nt of DNA (the human genome – the sum total of our DNA – is around 3 billion nt long). Interpreting these massive volumes of short reads is an ongoing challenge as sequencing technologies evolve. There are two popular analysis methods that process HTS reads: alignment-based and assembly-based approaches. The first uses a reference genome, whose DNA sequence is known from previous studies of the same or a closely related species. In this approach, reads are aligned to the reference genome through a process that searches for sequence similarities between the reads and the reference. The second approach is a data-driven method that does not assume similarity to any given genome. Instead, it reconstructs the genome represented by the DNA de novo (from scratch). This is a less biased approach that gives a truer representation of the genome, especially if there have been rearrangements compared to the reference genome sequence, or if no reference is available. The Birol lab has developed de novo assembly algorithms and downstream analysis tools and has applied them in a number of highly visible projects in human health and other fields. In the proposed work, the team will concentrate on alignment technologies as a way to support this highly successful assembly based analysis platform. The read alignment problem has been addressed several times, to match changes in read lengths and data volumes as HTS technology evolved. However, efficient and accurate alignment of reads to newly assembled genomes is an un-answered need. General purpose read alignment algorithms assume the target sequence to be composed of a small number of long stretches of sequence, essentially, chromosomes. The results of draft de novo assembly processes, in contrast, are typically in hundreds of thousands of pieces. This creates problems for general-purpose aligners, which we will address by developing an algorithm for this specific need. We will pay special attention to the scalability of our algorithm to accommodate the growing volume of data, and we will achieve this by building parallel processing algorithms similar to those used in Internet search engines, such as Google. When the genome of a new species is sequenced and assembled, one important task is to “annotate” its genes – i.e. mark where they are in the genome, and how they are structured. We also note an important gap in this area, as current alignment technologies were developed for previous generations of sequencing platforms, and have exceeded their limits to support data from new sequencing projects. (One such popular tool, exonerate, is still being heavily used, yet it is no longer being maintained by the developer lab.) We propose to build an alternative to these tools, and provide sustained support for the community. As the use of sequencing technologies further penetrates life sciences, there is an urgent need for high-quality computational tools to analyze large volumes of data in a timely manner. Development of the described alignment technologies will improve the efficiency and the accuracy of de novo assemblies and their annotation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2022
  • 负责人:
    Birol, Inanc
  • 依托单位:
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2021
  • 负责人:
    Birol, Inanc
  • 依托单位:
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2020
  • 负责人:
    Birol, Inanc
  • 依托单位:
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2019
  • 负责人:
    Birol, Inanc
  • 依托单位:
海外基金