课题基金 / 基金详情

AF: Medium: Collaborative Research: Sequential and Parallel Algorithms for Approximate Sequence Matching with Applications to Computational Biology

AF: Medium: Collaborative Research: Sequential and Parallel Algorithms for Approximate Sequence Matching with Applications to Computational Biology
AF:媒介:协作研究:近似序列匹配的顺序和并行算法及其在计算生物学中的应用
批准号:
1703489
负责人:
Sharma Thankachan
金额:
$29.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2021-06-30

项目摘要

项目成果

Sharma Thankachan的其他基金

相似基金

相关文献

中文摘要
翻译
序列匹配问题是基因组学领域的核心,无论是分析自然发生的序列,如基因组,还是分析测序仪器的数据。通常,在实践中,能够适应匹配区域内少量差异的方法就足够了。这种方法被描述为无对齐或近似序列匹配方法,通常依赖于启发式方法。该项目工作通过创建数学框架和解决具有可证明的高效运行时间保证的多个近似序列匹配问题来推进该领域。项目工作还支持由数学框架启发和支持的实用启发式的发展,开发用于在高性能并行计算机上解决大规模问题的并行方法,并研究这些方法对重要应用的影响。项目结果通过开放源代码软件提供,供实践者使用。这项研究的结果将纳入pi教授的课程,并通过书籍章节、教程和随附的幻灯片更广泛地传播。该项目将支持研究科学家和博士生进行跨学科培训,使他们进入专注于当前相关重要问题的生产性职业生涯。本科生参与计划通过课程项目。项目工作建立在无比对基因组比较方法的最新进展基础上,并利用高通量测序仪产生的数据的可控误差特征,以及它们所启用的许多生物信息学应用。项目目标包括开发一个健壮的算法框架,用于设计基于近似子串组成的新的无对齐方法,以及开发用于大型序列数据集之间的成对近似序列匹配的顺序和并行算法。目标是开发渐近优于基于二次对齐的方法的算法,并直接或通过依赖于底层理论的实际启发式的进一步发展获得良好的实际性能。所开发的技术将在诸如读取错误校正、基因组定位和组装等重要应用的背景下进一步研究。虽然是在计算生物学的背景下进行的,但其中一些方法可能适用于其他领域,如文本处理和信息检索。通过软件模块的发布和重要应用领域的项目工作,更广泛的研究社区将受到影响。
英文摘要
Sequence matching problems are central to the field of genomics, both in analyzing naturally occurring sequences such as genomes and in analyzing data from sequencing instruments. Often, methods that can accommodate a small number of differences within the matching regions suffice in practice. Such methods, described as alignment-free or approximate sequence matching methods, have typically relied on heuristics. This project work is advancing the field by creating a mathematical framework and solving multiple approximate sequence matching problems with provably efficient run-time guarantees. Project work is also supporting the development of practical heuristics inspired and supported by the mathematical framework, development of parallel methods for solving large-scale problems on high performance parallel computers, and studying the impact of these methods on important applications. Project results are made available through open source software for use by practitioners. Results from this research will be incorporated into courses taught by the PIs, and disseminated more broadly through book chapters and tutorials and accompanying slides. The project will support research scientist and Ph.D. students in interdisciplinary training for launching them into productive careers focused on important problems of current relevance. Undergraduate participation is planned through course projects.Project work builds upon recent progress in alignment-free genome comparison methods, and exploits the controlled error characteristics of data generated by high-throughput sequencers, and the many bioinformatics applications enabled by them. Project objectives include developing a robust algorithmic framework for designing newer alignment-free methods based on approximate substring composition, and developing sequential and parallel algorithms for pairwise approximate sequence matching among large sequence data sets. The goal is to develop algorithms that are asymptotically superior to quadratic alignment-based approaches, and achieve good practical performance either directly or through further development of practical heuristic that rely on the underlying theory. The developed techniques will be further investigated in the context of important applications such as read error correction, genome mapping, and assembly. Though conducted in the context of computational biology, some of the methods are potentially applicable to other areas such as text processing and information retrieval. Broader research community will be impacted through release of software modules and project work in important application areas.
期刊论文(17)
专著(0)
科研奖励(0)
会议论文
The Heaviest Induced Ancestors Problem: Better Data Structures and Applications
最严重的诱发祖先问题:更好的数据结构和应用程序
DOI: 10.1007/s00453-022-00955-7
发表时间: 2022
期刊: Algorithmica
影响因子: 1.1
作者: [Abedin, Paniz, Hooshmand, Sahar, Ganguly, Arnab, Thankachan, Sharma V.]
通讯作者: Thankachan, Sharma V.
On the Complexity of Recognizing Wheeler Graphs
论识别惠勒图的复杂性
DOI: 10.1007/s00453-021-00917-5
发表时间: 2022
期刊: Algorithmica
影响因子: 1.1
作者: [Gibney, Daniel, Thankachan, Sharma V.]
通讯作者: Thankachan, Sharma V.
A Linear-Space Data Structure for Range-LCP Queries in Poly-Logarithmic Time
多对数时间内范围LCP查询的线性空间数据结构
DOI: 10.1007/978-3-319-94776
发表时间: 2018
期刊: International Computing and Combinatorics Conference
影响因子: --
作者: [Abedin, P., Ganguly, A., Hon, W. K., Nekrich, Y., Sadakane, K., Shah, R., Thankachan, S. V.]
通讯作者: Thankachan, S. V.
The Heaviest Induced Ancestors Problem Revisited
重温最重的诱发祖先问题
DOI: 10.4230/lipics.cpm.2018.20
发表时间: 2018
期刊: {CPM} 2018
影响因子: --
作者: [Abedin, P., Hooshmand, S., Ganguly, A., Thankachan, S.V.]
通讯作者: Thankachan, S.V.
15
    REU Site: Algorithm Design --- Theory and Engineering
    • 批准号:
      2349179
    • 项目类别:
      Standard Grant
    • 资助金额:
      $46.24万
    • 财政年份:
      2024
    • 负责人:
      Sharma Thankachan
    • 依托单位:
    AF: Small: Theoretical Aspects of Repetition-Aware Text Compression and Indexing
    • 批准号:
      2315822
    • 项目类别:
      Standard Grant
    • 资助金额:
      $44.98万
    • 财政年份:
      2023
    • 负责人:
      Sharma Thankachan
    • 依托单位:
    CAREER: Algorithmic Aspects of Pan-genomic Data Modeling, Indexing and Querying
    • 批准号:
      2316691
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $79.53万
    • 财政年份:
      2023
    • 负责人:
      Sharma Thankachan
    • 依托单位:
    CAREER: Algorithmic Aspects of Pan-genomic Data Modeling, Indexing and Querying
    海外基金