课题基金 / 基金详情

Integrative Machine Learning Models for Discovery and Validation of Biological Knowledge

Integrative Machine Learning Models for Discovery and Validation of Biological Knowledge
用于发现和验证生物知识的综合机器学习模型
批准号:
RGPIN-2019-04696
负责人:
Rueda, Luis
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31

项目摘要

项目成果

Rueda, Luis的其他基金

相似基金

相关文献

中文摘要
翻译
下一代测序的出现彻底改变了基因组、转录组和外显子组的研究方式,通常会产生大量需要分析的数据。在这方面,机器学习已经提出了非常成功的方法来应用于生物现象的知识发现、检索和预测。尤其是深度学习,在生物、医学、数据和网络安全、文本挖掘和计算机视觉等领域的广泛应用中已被成功地应用于提取知识。然而,现有的数据和现有方法还存在一些挑战和局限性,以及一些有待克服的障碍,包括缺乏对不同变量的注释,缺乏大规模的和标记的训练数据,变量的多样性,不同数据集之间的格式差异,以及提取的知识缺乏样本特异性。*本研究计划旨在开发集成的机器学习系统,用于从多个变量、多个数据集和大量样本以及不同的生物学指标中提取相关的生物学知识,包括组学、文本和图形。综合模型将涉及不同类型疾病的多个数据集、多个组学和多个变量。在不同的子项目中,我们计划应用多模式、多任务和迁移(深度和浅层)机器学习方法,集成不同类型的数据。要开发的方法将整合不同的表征深度学习方案,这些方案利用不同形式的训练,如对手网络、卷积网络和循环网络,有或没有记忆。*综合机器学习方法的发展还没有出现在多组学数据中,因此,将文本数据与不同形式的分子测量相结合是开发新方法的一条很有前途的途径,然后这些方法也可以用于其他领域,例如数据安全、网络和添加剂制造,仅举几例。缺乏可靠的算法来整合和消除不一致数据的歧义是至关重要的,因为它们对于缺失的变量也是如此,因此,在这方面,综合的半监督方法是至关重要的。将要开发的方法将被其他研究人员使用,通过共享出版物和一个将部署在标准生物信息学和开放源码平台中的系统,从大数据集中发现生物知识。此外,预计在该研究计划下,将培训一名博士后研究员、三名博士生以及几名硕士和本科生,获得大数据分析以及工具和平台软件开发方面的关键技能。
英文摘要
The advent of next generation sequencing has revolutionized the way the genome, transcriptome and exome are studied, typically producing huge amounts of data to be analyzed. In this regard, machine learning has proposed a paramount of successful methods applied to knowledge discovery, retrieval and prediction of biological phenomena. Deep learning, in particular, has been successfully applied to extract knowledge in a wide range of applications in biology, medicine, data and network security, text mining, and computer vision, just to mention a few. There are, however, some challenges and limitations in the available data and the current approaches, as well as some obstacles yet to overcome, including lack of annotation of different variables, lack of large-scale and labelled training data, variety of variables, variations of formats across different datasets, and lack of sample-specificity in the knowledge extracted.******This research program aims to develop integrative machine learning systems used to extract relevant biological knowledge from multiple variables, multiple datasets with large numbers of samples, and different biological indicators including “-omics”, text and graphics from the literature. The integrative model will involve multiple datasets, multi-omics and multiple variables of different types of diseases. In the different sub-projects, we are planning to apply multi-modal, multi-task and transfer (deep and shallow) machine learning approaches that integrate different types of data. The approaches to be developed will integrate different schemes of representational deep leaning that utilize different forms of training such as adversary, convolutional and recurrent networks, with or without memory. ******Development of integrative machine learning approaches have not emerged significantly in multi-omics data, and hence, incorporating textual data, integrated with molecular measurements of different forms is a promising avenue for developing novel approaches, which can be then used in other fields as well, such as data security, networking and additive manufacturing, just to mention a few. The lack of reliable algorithms for integrating and disambiguating inconsistent data are crucial, as they are for missing variables thus, integrative semi-supervised approaches are crucial in this regard. The methods to be developed will be used by other researchers in discovering biological knowledge from large datasets via sharing publications and a system that will be deployed in standard bioinformatics and open source platforms. In addition, it is expected that under this research program one postdoctoral fellow, three PhD students, and several Master's and undergraduate students will be trained, gaining key skills in big data analytics and software development of tools and platforms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Integrative Machine Learning Models for Discovery and Validation of Biological Knowledge
  • 批准号:
    RGPIN-2019-04696
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2022
  • 负责人:
    Rueda, Luis
  • 依托单位:
Integrative Machine Learning Models for Discovery and Validation of Biological Knowledge
  • 批准号:
    RGPIN-2019-04696
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2021
  • 负责人:
    Rueda, Luis
  • 依托单位:
NSERC I2I Phase Ia: An Intelligent Framework for Social Engineering Cyber Security Training
  • 批准号:
    567660-2021
  • 项目类别:
    Idea to Innovation
  • 资助金额:
    $9.11万
  • 财政年份:
    2021
  • 负责人:
    Rueda, Luis
  • 依托单位:
Market Assessment of an intelligent framework for social engineering cyber security training
  • 批准号:
    556923-2020
  • 项目类别:
    Idea to Innovation
  • 资助金额:
    $1.09万
  • 财政年份:
    2020
  • 负责人:
    Rueda, Luis
  • 依托单位:
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位: