课题基金 / 基金详情

Read-to-contig alignments for de novo genome assembly and annotation

Read-to-contig alignments for de novo genome assembly and annotation
用于从头基因组组装和注释的读取到重叠群比对
批准号:
RGPIN-2014-05112
负责人:
Birol, Inanc
金额:
$2.57万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2017
资助国家:
加拿大
项目状态:
已结题
起止时间:
2017-01-01 至 2018-12-31

项目摘要

项目成果

Birol, Inanc的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The proposed research is about building computational technologies to analyze DNA. DNA is composed of sequences of four possible nucleotides (nt): A, C, G and T. The last decade witnessed a revolution in technologies that “read” DNA sequences, with applications in many areas of life sciences.Data from high throughput sequencing (HTS) platforms reach hundreds of millions of “reads”, where each read represents 75-300 nt of DNA (the human genome – the sum total of our DNA – is around 3 billion nt long). Interpreting these massive volumes of short reads is an ongoing challenge as sequencing technologies evolve.There are two popular analysis methods that process HTS reads: alignment-based and assembly-based approaches. The first uses a reference genome, whose DNA sequence is known from previous studies of the same or a closely related species. In this approach, reads are aligned to the reference genome through a process that searches for sequence similarities between the reads and the reference. The second approach is a data-driven method that does not assume similarity to any given genome. Instead, it reconstructs the genome represented by the DNA de novo (from scratch). This is a less biased approach that gives a truer representation of the genome, especially if there have been rearrangements compared to the reference genome sequence, or if no reference is available.The Birol lab has developed de novo assembly algorithms and downstream analysis tools and has applied them in a number of highly visible projects in human health and other fields. In the proposed work, the team will concentrate on alignment technologies as a way to support this highly successful assembly based analysis platform.The read alignment problem has been addressed several times, to match changes in read lengths and data volumes as HTS technology evolved. However, efficient and accurate alignment of reads to newly assembled genomes is an un-answered need. General purpose read alignment algorithms assume the target sequence to be composed of a small number of long stretches of sequence, essentially, chromosomes. The results of draft de novo assembly processes, in contrast, are typically in hundreds of thousands of pieces. This creates problems for general-purpose aligners, which we will address by developing an algorithm for this specific need. We will pay special attention to the scalability of our algorithm to accommodate the growing volume of data, and we will achieve this by building parallel processing algorithms similar to those used in Internet search engines, such as Google.When the genome of a new species is sequenced and assembled, one important task is to “annotate” its genes – i.e. mark where they are in the genome, and how they are structured. We also note an important gap in this area, as current alignment technologies were developed for previous generations of sequencing platforms, and have exceeded their limits to support data from new sequencing projects. (One such popular tool, exonerate, is still being heavily used, yet it is no longer being maintained by the developer lab.) We propose to build an alternative to these tools, and provide sustained support for the community.As the use of sequencing technologies further penetrates life sciences, there is an urgent need for high-quality computational tools to analyze large volumes of data in a timely manner. Development of the described alignment technologies will improve the efficiency and the accuracy of de novo assemblies and their annotation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2022
  • 负责人:
    Birol, Inanc
  • 依托单位:
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2021
  • 负责人:
    Birol, Inanc
  • 依托单位:
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2020
  • 负责人:
    Birol, Inanc
  • 依托单位:
Novel Data Structures And Scalable Algorithms For High Throughput Bioinformatics
  • 批准号:
    RGPIN-2019-06640
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2019
  • 负责人:
    Birol, Inanc
  • 依托单位:
海外基金