课题基金 / 基金详情

项目摘要

项目成果

SANGUTHEVAR RAJASEKARAN的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):多个基因组计划已经产生了大量的DNA、RNA和蛋白质序列数据。虽然计算模式搜索技术,如BLAST(用于测量序列相似性的程序)已经实现了重大发现,如蛋白质的模块化,但可以从生物序列数据中获得更多的信息。然而,由于模式的长度和复杂性,我们受到用于模体发现的计算算法的限制。简单基序搜索(SMS),种植基序搜索(PMS)和编辑距离基序搜索(EMS)是三个主要的范式,以前已被用于识别短功能肽基序,转录调控元件,复合调控模式,DNA基序,蛋白质家族之间的相似性,等我们的小组一直在开发这些问题的算法。 现有的模体搜索的模式搜索算法有两个主要的缺点:1)近似算法并不总是识别正确的模式,但有一个优点,他们可以用来寻找短的和相对较大的模式在大数据集,如基因组。2)相比之下,精确算法总是识别正确的模式,但不能用于识别大型数据集中的复杂数据模式。为了从基因组数据中提取更复杂的模式,我们需要精确的算法,可以用来分析基因组的复杂模式与合理的计算资源。精确的算法目前是有限的,因为这些算法的运行时间是指数依赖于所涉及的参数。例如,目前最知名的PMS和EMS算法(在PC上)预计长度为27的模式需要一个多月,长度为31的模式需要5.67年以上。在这个项目中,我们建议开发下一代SMS,PMS和EMS算法,这些算法可以使用更少的计算时间和内存在更大的数据集中识别更复杂的模式。我们还建议开发一个基于Web的系统,将PMS和EMS算法。开发的所有算法和数据的开源版本将提供给用户。此外,网络系统将支持在线处理涉及PMS和EMS解决方案的查询。
英文摘要
DESCRIPTION (provided by applicant): Multiple genome projects have generated large volumes of DNA, RNA and protein sequence data. While computational pattern searching techniques such as BLAST (a program for measuring sequence similarities) have enabled major discoveries such as the modularity of proteins, much more information can be gained from biological sequence data. However, due to the length and complexity of the patterns, we are limited by the computational algorithms used for motif discovery. Simple Motif Search (SMS), Planted Motif Search (PMS) and Edit-distance Motif Search (EMS) are the three principal paradigms that have been previously used for identifying short functional peptide motifs, transcriptional regulatory elements, composite regulatory patterns, DNA motifs, similarity between families of proteins, etc. Our group has been instrumental in developing algorithms for these problems. Existing pattern-search algorithms for motif search have two major shortcomings: 1) Approximate algorithms do not always identify the correct pattern, but have the advantage that they can be used to look for short and relatively large patterns in large data sets such as genomes. 2) In contrast, an exact algorithm always identifies the correct pattern, but cannot be used to identify complex data patterns in large datasets. To extract more sophisticated patterns from genomic data we need exact algorithms that can be used to analyze genomes for complex patterns with reasonable computational resources. Exact algorithms are currently limited because the run times of these algorithms are exponentially dependent on the parameters involved. For example, the currently best known algorithms for PMS and EMS (on a PC) are expected to take more than a month for patterns of length 27 and more than 5.67 years for patterns of length 31. In this project, we propose to develop the next generation SMS, PMS and EMS algorithms that identify more complex patterns in larger datasets using less computation time and memory. We also propose to develop a web based system that will incorporate PMS and EMS algorithms. Open source versions of all the algorithms and data developed will be made available to users. In addition, the web system will support online processing of queries that involve the solution of PMS and EMS.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Efficient Algorithms for Motif Search
  • 批准号:
    8142235
  • 项目类别:
  • 资助金额:
    $35.52万
  • 财政年份:
    2010
  • 负责人:
    SANGUTHEVAR RAJASEKARAN
  • 依托单位:
Efficient Algorithms for Motif Search
  • 批准号:
    7878215
  • 项目类别:
  • 资助金额:
    $39.7万
  • 财政年份:
    2010
  • 负责人:
    SANGUTHEVAR RAJASEKARAN
  • 依托单位:
海外基金