课题基金 / 基金详情

Efficient Representation and Manipulation of Large-Scale Biological Sequence Data

Efficient Representation and Manipulation of Large-Scale Biological Sequence Data
大规模生物序列数据的高效表示和操作
批准号:
0430853
负责人:
Srinivas Aluru
金额:
$44.05万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-09-01 至 2008-08-31

项目摘要

项目成果

Srinivas Aluru的其他基金

相似基金

相关文献

中文摘要
翻译
生物分子序列的存储和访问以确定序列同源性是当前生物信息学和计算生物学革命的核心。除了搜索工具外,一些重要应用程序使用的大量生物数据突显了开发高效的核外算法的必要性。该项目的目标是设计磁盘驻留序列数据的存储结构、算法技术和软件,并将其应用于计算生物学中的重要应用。为了实现这一目标,我们使用了一个三管齐下的策略:首先,与领域专家合作确定的应用需求被用来设计序列数据的基本存储结构。这项研究包括为众所周知的核内数据结构开发高效的核外算法,以及设计适合目标应用的新数据结构。其次,正在开发对驻留在磁盘上的序列数据进行高效查询的算法。最后,将开发的核心外技术与计算基因组学中的应用软件集成,如EST聚类法和片段拼接。其目标是开发更快的算法,减少过高的主存需求,或酌情解决更大的问题实例。研究结果将以软件库的形式提供给计算机科学家,并以应用软件的形式提供给分子生物学家。正在努力将这项研究的结果纳入分子生物学家使用的流行工具中。该项目的跨学科性质为研究生提供了独特的培训机会。
英文摘要
ABSTRACTStorage of biomolecular sequences, and accessing them to determine sequence homologies is central to the current revolution in bioinformatics and computational biology. Besides search tools, the large size of biological data used by some important applications underscores the need for developing efficient out-of-core algorithms. The goal of the project is to design storage structures, algorithmic techniques, and software for disk-resident sequence data, and apply it to important applications in computational biology. To achieve thisgoal, a three-pronged strategy is used: Firstly, application requirements identified in collaboration with domain experts are being used to design fundamental storage structures for sequence data. This research spans the development of efficient out-of-core algorithms for well-known in-core data structures and also the design of new data structures suitable for targeted applications. Secondly, efficient algorithms for queries on disk-resident sequence data are being developed. Finally, the out-of-core techniques developed are integrated with application software in computational genomics such as EST clustering and fragment assembly. The goal is to develop faster algorithms, reduce the exorbitant main-memory requirements, or enable solution of larger problem instances, as appropriate.The results of the research will be made accessible to computer scientistsin the form of software libraries and molecular biologists in the form of application software. Efforts are being made to integrate the results of this research into popular tools used by molecular biologists. The interdisciplinary nature of the project is providing unique training opportunities for graduate students.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A scalable integrated multi-modal single cell analysis framework for gene regulatory and cell-cell interaction networks
  • 批准号:
    2233887
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $54.58万
  • 财政年份:
    2023
  • 负责人:
    Srinivas Aluru
  • 依托单位:
BD Hubs: Collaborative Proposal: SOUTH:The South Big Data Innovation Hub
  • 批准号:
    1916589
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $203.16万
  • 财政年份:
    2019
  • 负责人:
    Srinivas Aluru
  • 依托单位:
AF: Small: Algorithmic Techniques for High-throughput Analysis of Long Reads
  • 批准号:
    1816027
  • 项目类别:
    Standard Grant
  • 资助金额:
    $42.5万
  • 财政年份:
    2018
  • 负责人:
    Srinivas Aluru
  • 依托单位:
EAGER: A Framework for Learning Graph Algorithms with Applications to Social and Gene Networks
  • 批准号:
    1841351
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2018
  • 负责人:
    Srinivas Aluru
  • 依托单位:
海外基金