课题基金 / 基金详情

Novel domain-specific languages and compiler optimization methods for computational biology

Novel domain-specific languages and compiler optimization methods for computational biology
计算生物学的新颖的特定领域语言和编译器优化方法
批准号:
RGPIN-2019-04973
负责人:
Numanagic, Ibrahim
金额:
$2.04万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31

项目摘要

项目成果

Numanagic, Ibrahim的其他基金

相似基金

相关文献

中文摘要
翻译
动机:下一代测序(NGS)实验产生的海量数据需要开发高效的计算方法,因为计算方面是目前NGS管道的最大瓶颈。然而,许多有希望的方法无法应对当前和新兴的NGS技术的规模,而且太难使用和复制。这个问题的根本原因在于广泛使用的通用开发环境无法有效地表达和优化生物数据工作流。用户被迫要么使用高级但速度慢的语言,如Python,要么使用低级语言,如C语言,这些语言可以产生高效的工具,但需要花费大量的时间和可维护性成本。*方法:一种专门为生物和测序数据量身定做的领域特定语言(DSL)和相关的编译器优化技术将为实验新的计算算法提供灵活性、简单性和模块化,同时生成高性能代码并使程序从资源受限的体系结构移植到最大的超级计算机上。*目标:我们提出了一种新的领域特定语言和相关编译器Seq,它能够快速、轻松地开发高性能测序流水线。为了实现这一目标,我们将:*(I)设计一种编程语言和编译器,使诸如Python之类的高级语言易于开发,同时提供诸如C(目标1)之类的低级语言的原始性能;*(Ii)探索基因组工作流中的数据访问模式,并设计能够在编译器级别利用这些模式的方法,以跨各种计算环境(例如多核CPU、GPU和手持设备)进行自动低级优化(目标2);以及*(三)提供将SEQ轻松整合到流行的生物信息学和科学环境中的手段,并为NGS数据开发一个经过策划的算法原语库(目标3)。*近期目标是开发一款能够在常见架构上高效处理各种NGS数据的DSL。长期目标是建立一个全面和广泛使用的基础设施,允许快速和轻松地开发生物数据的方法。由于计算生物学HQP在加拿大的需求量很大,这项提议的主要目标之一是在五年的时间里培训HQP。*影响:我们设想我们的DSL将通过使研究人员能够以更自然的方式表达他们的想法并允许他们使用最佳算法方法来显著促进加拿大的基因组学和健康研究。此外,我们预计我们的DSL将通过提供大量的时间和成本节省来帮助加拿大的大型科学卫生项目。我们还预计SEQ将成为广泛使用的生物信息学工具的关键组成部分。最后,我们期待通过该计划培训的HQP将为加拿大的知识型经济做出贡献。**
英文摘要
Motivation: The vast scale of data generated by next-generation sequencing (NGS) experiments necessitates the development of efficient computational methods, as the computational aspect is currently the biggest bottleneck of NGS pipelines. However, many promising methods cannot handle the scale of current and emerging NGS technologies and are too hard to use and replicate. Root cause of this problem lies in widely used general-purpose development environments that cannot efficiently express and optimize biological data workflows. Users are forced to use either high-level but slow languages such as Python, or low-level languages such as C that produce efficient tools but at a significant time and maintainability costs. ***Approach: A domain-specific language (DSL) and associated suite of compiler optimization techniques specifically tailored for biological and sequencing data would provide flexibility, simplicity and modularity for experimenting with new computational algorithms, while generating high-performance code and making the programs portable from resource-constrained architectures to the biggest supercomputers.***Objectives: We propose a novel DSL and associated compiler named Seq that enables rapid and easy development of high-performance sequencing pipelines. To achieve this, we will: ***(i) design a programming language and a compiler that allows ease of development of high-level languages such as Python, while providing raw performance of low-level languages such as C (Objective 1);***(ii) explore data access patterns in genomic workflows and devise methods that can exploit these patterns at the compiler level for automatic low-level optimizations across various computational environments, such as multicore CPUs, GPUs and handheld devices (Objective 2); and ***(iii) provide means to easily integrate Seq into popular bioinformatics and scientific environments and develop a curated library of algorithmic primitives for NGS data (Objective 3). ***The short-term goal is to develop a DSL that can efficiently handle various kinds of NGS data on common architectures. The long-term goal is to build a comprehensive and widely used infrastructure that allows rapid and easy method development for biological data. As computational biology HQP are in high demand in Canada, one of the key goals of this proposal is to train HQP over the course of five years.***Impact: We envision our DSL to significantly boost Canadian genomics and health research by enabling researchers to express their ideas in a more natural way and by allowing them to use the best algorithmic methods for the job. Furthermore, we expect our DSL to aid large-scale scientific Canadian health projects by providing huge time and cost savings. We also anticipate Seq to become a key building block in the wide specter of widely used bioinformatics tools. Finally, we expect that HQP trained by this program will contribute to the Canadian knowledge-based economy.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data Science
  • 批准号:
    CRC-2018-00136
  • 项目类别:
    Canada Research Chairs
  • 资助金额:
    $8.74万
  • 财政年份:
    2022
  • 负责人:
    Numanagic, Ibrahim
  • 依托单位:
Novel domain-specific languages and compiler optimization methods for computational biology
  • 批准号:
    RGPIN-2019-04973
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2022
  • 负责人:
    Numanagic, Ibrahim
  • 依托单位:
Data Science
  • 批准号:
    CRC-2018-00136
  • 项目类别:
    Canada Research Chairs
  • 资助金额:
    $8.74万
  • 财政年份:
    2021
  • 负责人:
    Numanagic, Ibrahim
  • 依托单位:
Novel domain-specific languages and compiler optimization methods for computational biology
  • 批准号:
    RGPIN-2019-04973
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2021
  • 负责人:
    Numanagic, Ibrahim
  • 依托单位:
国内基金
海外基金
Domain理论中几类T0拓扑空间的幂构造研究
  • 批准号:
    2026JJ81209
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    袁珍珠
  • 依托单位:
RB-domain函数空间的相关研究
  • 批准号:
    2026JJ60113
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    栾伟
  • 依托单位:
RIPK3蛋白及其RHIM结构域在脓毒症早期炎症反应和脏器损伤中的作用和机制研究
  • 批准号:
    82372167
  • 项目类别:
    面上项目
  • 资助金额:
    48.00万元
  • 批准年份:
    2023
  • 负责人:
    江继宏
  • 依托单位:
拟连续domain范畴的若干问题研究
  • 批准号:
    12301583
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2023
  • 负责人:
    栾伟
  • 依托单位: