课题基金 / 基金详情

Machine learning-based alignment-free methodology for complete genome analysis

Machine learning-based alignment-free methodology for complete genome analysis
基于机器学习的免比对方法,用于完整基因组分析
批准号:
RGPIN-2022-03547
负责人:
Randhawa, Gurjit
金额:
$1.82万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Randhawa, Gurjit的其他基金

相似基金

相关文献

中文摘要
翻译
序列分类是根据生物体基因组序列的差异和相似性对生物体进行识别、命名和分组的科学实践。序列分类的问题是非常重要的,考虑到我们这个星球上估计有870万(± 130万)种物种,到目前为止只有大约150万种不同的真核生物被编目。地球上86%的现存物种和91%的海洋物种仍然未分类。由于数据集的规模和复杂性,用于分类目的的序列比较和分析问题仍然具有挑战性。 作为数据科学助理教授,我的工作包括设计和开发基于机器学习的模型,以分析自然科学和工程各个领域产生的复杂数据结构。在接下来的五年里,我将考虑两个主要的应用主题:1)无比对的基因组数据分析和2)使用进化计算的模型优化。在主题1中,我将开发一个开源的、基于Web的、超快的、可扩展的、基于机器学习的无障碍工具,用于实时分析和准确分类基因组序列。在头两年将开发一个原型平台,以支持病毒基因组的分类。在接下来的两年内,将增加对细菌基因组和宏基因组数据的支持。作为另一个应用程序,我将基准的测序信息的最小百分比,这是一个必须准确的分类。对于任何特定物种,将使用过滤技术以受控方式选择序列特征的子集,以找到基因组签名(沿基因组沿着的物种特异性模式)长度的最小阈值。作为另一个应用,我将探索基因组签名的空间,以建立形成基因组签名和影响基因组完整性的不同潜在机制之间的定量关系。特别是,将研究环境诱变剂对序列组成的影响。这可能会回答多个有趣的问题,例如基因组签名中有多少信息是因为进化?暴露于环境诱变剂能提供多少信息?在主题2中,我将研究如何使用进化计算优化机器学习模型的训练过程。作为该主题下的另一个应用,我计划基于成对距离定义物种特异性的定量基因组特征图谱。然后,遗传算法将被应用于发展分类规则集的基础上,这些签名配置文件,后来被用来预测新的未知序列的标签。如果成功的话,这些新颖的时间效率模型的使用可以扩展到解决其他学科的分类问题。
英文摘要
Sequence classification is the scientific practice of identifying, naming, and grouping organisms based on differences and similarities in their genomic sequences. The problem of sequence classification is of immense importance considering that out of estimated 8.7 million (±1.3 million) species on our planet, only around 1.5 million distinct eukaryotes have been catalogued so far. This leaves us with 86% of existing species on Earth and 91% of marine species still unclassified. Due to the magnitude and complexity of the datasets, the problem of sequence comparison and analysis for the purpose of classification remains challenging.  As an Assistant Professor in Data Science, my work consists of designing and developing machine learning-based models to analyze complex data structures arising from various areas in the Natural Sciences and Engineering. I will consider two main application themes over the next five years: 1) Alignment-free genomic data analysis and 2) Model optimization using evolutionary computing. In Theme 1, I will develop an open-source, web-based, ultra-fast, and scalable machine learning-based alignment-free tool for real-time analysis and accurate classification of genomic sequences. A prototype platform will be developed to support the classification of viral genomes in the first two years. In the following two years, support for bacterial genomes and metagenomic data will be added. As another application, I will benchmark the minimum percentage of sequencing information that is a must for accurate classification. For any particular species, a subset of sequence features will be selected in a controlled manner using filtering techniques to find the minimum threshold on the length of a genomic signature (a species-specific pattern that is pervasive along the genome). As another application, I will explore the space of the genomic signature to establish a quantitative relationship between different underlying mechanisms that shape genomic signature and affect genomic integrity. In particular, the effect of environmental mutagens on the sequence composition will be studied. This may answer multiple interesting questions such as how much information in the genomic signature is because of evolution?, How much information is contributed by exposure to environmental mutagens? etc. In Theme 2, I will conduct research on optimizing the training process of machine learning models using evolutionary computing. As another application under this theme, I plan to define species-specific, quantitative genomic signature profiles based on pairwise distances. Genetic algorithms will then be applied to evolve classification rule sets based on these signature profiles and later be used to predict the labels of new unknown sequences. If successful, the use of these novel time-efficient models can then be extended to address the classification problems from other disciplines.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Machine learning-based alignment-free methodology for complete genome analysis
  • 批准号:
    DGECR-2022-00370
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2022
  • 负责人:
    Randhawa, Gurjit
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: