课题基金 / 基金详情

Algorithms in comparative genomics and ancestor reconstructions

Algorithms in comparative genomics and ancestor reconstructions
比较基因组学和祖先重建的算法
批准号:
RGPIN-2014-05729
负责人:
Bergeron, Anne
金额:
$1.89万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Bergeron, Anne的其他基金

相似基金

相关文献

中文摘要
翻译
在这个项目中,我们将开发组合模型和算法来研究计算生物学中的两个问题,这两个问题与重建祖先序列的相关性有关,以了解某些结构是如何进化的。第一个问题是哺乳动物中可选转录本的比较和预测,第二个问题是噬菌体中模块重组的重建。 在2000年对哺乳动物基因组进行测序之前,大多数科学家认为,考虑到哺乳动物的复杂性,它们的基因应该比昆虫、蠕虫或植物多得多。事实证明,人类和蠕虫的基因数量大致相同,但往往比植物的基因少。ENCODE项目(2207)的结果部分解决了这一明显的悖论,该项目的结果表明,大多数哺乳动物的基因能够使用一种被称为“替代转录”的机制来制造多达数十种不同的蛋白质。我们最近开发了一个组合模型,允许比较不同物种中相似基因产生的蛋白质组。我们展示了可以使用包含注释的树和字符串来模拟这些可选转录集的演变。我们现在的目标是将这些描述性工具转变为预测工具,以利用现在可以常规获取的记录数据的身份和丰富的优势。 这项研究的一个具体成果将是生命科学家可以用来预测跨物种替代转录的软件:这可能会产生重要的影响,因为它允许将从模型生物--如大鼠或小鼠--获得的结果转移到人类身上。一个更根本的结果是开发出能够处理基因结构和进化复杂性的抽象工具。 我们程序的第二部分专注于感染细菌的病毒,这种病毒被称为噬菌体,它将重组作为一个激进的发明过程:两个噬菌体基因组可能交换没有可检测到的相似性但具有相同生物功能的遗传序列。当相互比较时,得到的基因组共享相同的序列,其中夹杂着非常不同的序列。这种现象是在上个世纪观察到的,并导致了模块化理论的产生,该理论假设噬菌体是具有相同生物功能的模块的集合,这些模块可能带有不同的工具。给出一组噬菌体,我们的目标是重建这些重组的历史。第一步是确定集合中可能的重组:这是通过使用相似的序列和注释的模块作为锚,对所有基因组的模块进行“比对”来完成的。在第一篇论文中,我们能够部分比对几十个基因组,即使它们的大部分序列缺乏任何可检测到的相似性或是注释的假设蛋白质。我们现在有信心,我们可以自动化这一过程,并将其扩展到大多数主要的噬菌体家族,前提是它们的基因组组织是共线性的。这些比对将提供重组组合研究所需的数据集。 噬菌体现在被广泛认为是最多样化和数量最多的“生命”形式之一。从人类的角度来看,它们可以成为我们的朋友,通过传播使细菌对抗生素产生抗药性的遗传物质来杀死有害细菌或敌人。通过在这个项目中开发的计算模型和软件,我们的目标是进一步了解它们的多样性和它们的进化机制。
英文摘要
In this program, we will develop combinatorial models and algorithms for the study of two problems in computational biology that are related by the relevance of reconstructing ancestral sequences in order to understand how certain structures have evolved. The first problem is the comparison and prediction of alternative transcripts in mammals, and the second is the reconstruction of modular recombinations in phages. Before the sequencing of mammal genomes, starting with the human genome in 2000, most scientists believed that, given their complexity, mammals should have much more genes than insects, worms, or plants. It turned out that human beings and worms have about the same number of genes, and often less genes than plants. The key to this apparent paradox was partially solved by the results of the ENCODE project (2207), that showed that most mammal genes are able to manufacture up to dozens of different proteins using a mechanism called 'alternative transcription'. We recently developed a combinatorial model that allows the comparison of sets of proteins produced by similar genes in different species. We showed that it is possible to model the evolution of these sets of alternative transcripts using trees and strings containing annotations. We now aim to transform these descriptive tools into predictive tools for transcript discovery, taking advantage of the identity and abundance of transcript data that can now be routinely acquired. A tangible output of this research will be software that can be used by life scientists to predict alternative transcription across species: this can have important repercussions by allowing the transfer of results obtained with model organisms -- such as rats or mice -- to human beings. A more fundamental outcome is the development of abstract tools that can deal with the complexity of gene structure and evolution. The second part of our program focuses on viruses that infect bacteria, called phages, that use recombination as a radical invention process: two phage genomes may exchange genetic sequences that have no detectable similarity but have the same biological function. When compared to one another, the resulting genomes share identical sequences interspersed with very different sequences. This phenomenon was observed in the last century and gave rise to the modular theory which postulates that phages are assemblages of modules that carry identical biological function with possibly different tools. Given a set of phages, our goal is to reconstruct the history of these recombinations. The first step is to identify putative recombinations within the set: this is done by 'aligning' the modules of all genomes, using similar sequences and annotated modules as anchors. In a first paper, we were able to partially align a few dozen genomes, even if large parts of their sequences lack any detectable similarity or are annotated hypothetical proteins. We are now confident that we can automate the procedure and extend it to most major families of phages, provided that their genome organization is co-linear. These alignments will provide the datasets needed for a combinatorial study of recombinations. Phages are now widely recognized as one of the most diverse and numerous form of 'life'. From a human point of view, they can be our friends, by killing harmful bacteria, or foes, by propagating genetic material that make bacteria resistant to antibiotics. With the computational models and software developed within this program, we aim to further understand their diversity and their mechanisms of evolution.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms in comparative genomics and ancestor reconstructions
  • 批准号:
    RGPIN-2014-05729
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2021
  • 负责人:
    Bergeron, Anne
  • 依托单位:
Algorithms in comparative genomics and ancestor reconstructions
  • 批准号:
    RGPIN-2014-05729
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2019
  • 负责人:
    Bergeron, Anne
  • 依托单位:
Algorithms in comparative genomics and ancestor reconstructions
  • 批准号:
    RGPIN-2014-05729
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2016
  • 负责人:
    Bergeron, Anne
  • 依托单位:
Algorithms in comparative genomics and ancestor reconstructions
  • 批准号:
    RGPIN-2014-05729
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2015
  • 负责人:
    Bergeron, Anne
  • 依托单位:
国内基金
海外基金
优化基因组策略搜寻中国藏族内耳畸形的致病基因及其致聋机制研究
  • 批准号:
    31071099
  • 项目类别:
    面上项目
  • 资助金额:
    40.0万元
  • 批准年份:
    2010
  • 负责人:
    戴朴
  • 依托单位: