课题基金 / 基金详情

Compact Data Structures for Computational Genomics

Compact Data Structures for Computational Genomics
计算基因组学的紧凑数据结构
批准号:
RGPIN-2020-07185
负责人:
Gagie, Travis
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Gagie, Travis的其他基金

相似基金

相关文献

中文摘要
翻译
最大的基因组数据库现在占用了数百TB的未压缩空间。它们对研究人员和医生来说是无价的资源,但对计算机科学家来说是一个挑战,因为标准工具在这样的规模下变得不切实际。我们可以非常好地压缩这些数据库,因为它们是重复的,但我们仍然必须修改我们的工具,以使用它们的压缩表示法。例如,图谱组装是生物信息学中的一项基本任务:为我们已经组装的参考基因组编制索引,然后,对于新基因组的每一次读取,快速找到与其最匹配的参考基因组的子串。然而,即使基因组与参考基因组的99.9%相同,基因组也是如此之大,0.1%的差异足以导致数千次读取保持未绘制的图谱。它有助于为代表许多基因组的“泛基因组”参考建立索引,但流行的读取映射工具无法处理包含多个基因组的数据库,因为它们没有利用基因组的相似性。我们最近设计了一种名为r-index的工具,它可以使用合理的存储空间为数百或数千个人类基因组编制索引。我们已经发布了一个基本的实际实施,它将作为这个项目的起点。然而,向r-index添加新的功能比向当前的读取映射器添加它们更困难,因为现在我们需要压缩甚至小的辅助数据结构。首先,我们将扩展r索引以找到最大的精确匹配,并将其与最近存储泛基因组作为变异图的软件相结合。目前,该软件试图为图形编制索引,但它也可以为基因组数据库编制索引,并存储从每个基因组到图形中路径的映射。通过这种方式,给定一个读取,我们首先将其映射到基因组中的匹配子字符串,然后将这些映射到变异图中的路径。其次,我们将扩展r-index,以便当基因组数据库中有许多匹配的读取都映射到图中的同一路径时,它将只返回汇总统计信息,而不是单独报告它们。这样的查询本质上是高度重复的集合上的文档列表和文档计数,这是我以前研究过的。第三,我们将进一步研究Wheeler图,这是一种基于Burrow-Wheeler变换的数据结构框架。这样的数据结构包括作为流行的读映射器基础的FM索引;De Bruijn图的一些紧凑表示;以及变化图索引。我们的框架已经启发了一种进一步压缩r索引和压缩读数集索引的方法。我们将致力于将扩展r索引的实际实现集成到商业和研究测序管道中。我相信这个项目将帮助计算机科学家应对基因组数据洪流带来的一些挑战,从而帮助生物学家和医生实现其潜力。
英文摘要
The largest genomic databases now occupy hundreds of terabytes uncompressed. They are an invaluable resource for researchers and physicians but a challenge for computer scientists because standard tools become impractical at such scales. We can compress these databases extremely well because they are repetitive, but we must still modify our tools to work with their compressed representations. For example, mapping assembly is a basic task in bioinformatics: indexing a reference genome we have already assembled and then, for each read of a new genome, quickly finding the substrings of the reference it matches most closely. Even when a genome is 99.9% the same as the reference, though, genomes are so big that the 0.1% difference is enough to cause many thousands of reads to remain unmapped. It helps significantly to index a "pan-genome" reference representing many genomes but popular tools for read-mapping cannot handle databases containing more than a few genomes, because they do not take advantage of the genomes' similarity. We recently devised a tool called the r-index that can index hundreds or thousands of human genomes using reasonable memory space. We have released a basic practical implementation, which will serve as a starting point for this project. Adding new functionalities to the r-index is more difficult than adding them to current read-mappers, however, because now we need to compress even small auxiliary data structures. First, we will extend the r-index to find maximal exact matches and combine it with recent software that stores a pan-genome as a variation graph. At the moment that software tries to index a graph but it can just as well index a genomic database and store a mapping from each genome to a path in the graph. This way, given a read, we first map it to matching substrings in the genomes and then map those to paths in the variation graph. Second, we will extend the r-index such that, when there are many matches for a read in our genomic database that all map to the same path in the graph, it will return only summary statistics instead of reporting them individually. Such queries are essentially document listing and document counting over highly repetitive collections, something I have previously studied. Third, we will investigate further Wheeler graphs, a framework for data structures based on the Burrows-Wheeler Transform. Such data structures include the FM-index underlying popular read-mappers; some compact representations of de Bruijn graphs; and variation-graph indexes. Our framework has already inspired a method for further compressing the r-index and for compressed indexing of readsets. We will work to integrate practical implementations of an extended r-index into commercial and research sequencing pipelines. I am confident this project will help computer scientists meet some of the challenges posed by the flood of genomic data, and thus help biologists and physicians realize its potential.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Compact Data Structures for Computational Genomics
  • 批准号:
    RGPIN-2020-07185
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2021
  • 负责人:
    Gagie, Travis
  • 依托单位:
Compact Data Structures for Computational Genomics
  • 批准号:
    RGPIN-2020-07185
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2020
  • 负责人:
    Gagie, Travis
  • 依托单位:
Compact Data Structures for Computational Genomics
  • 批准号:
    DGECR-2020-00311
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2020
  • 负责人:
    Gagie, Travis
  • 依托单位:
PGSA
  • 批准号:
    231836-2000
  • 项目类别:
    Postgraduate Scholarships
  • 资助金额:
    $1.26万
  • 财政年份:
    2001
  • 负责人:
    Gagie, Travis
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: