课题基金 / 基金详情

III: Small: RUI: Efficient Search, Comparison, and Annotation for Biological Sequences

III: Small: RUI: Efficient Search, Comparison, and Annotation for Biological Sequences
III:小:RUI:生物序列的高效搜索、比较和注释
批准号:
1528027
负责人:
Abdullah Arslan
金额:
$7.68万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-01 至 2018-07-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目旨在开发快速搜索、比较和注释蛋白质和RNA序列的算法。正则表达式匹配通常在基于unix的系统中用于搜索文本,在PROSITE网站中用于搜索蛋白质序列中的模式。上下文无关语法匹配用于RNA序列的搜索。使用正则表达式和上下文无关语法描述基序来注释生物序列(DNA,蛋白质和RNA序列)是许多公共数据库,网站和软件工具(例如PROSITE, Locomotif)中可用的重要应用程序。该项目的结果将有助于开发模式匹配和解析应用程序的程序员,这些应用程序将受益于快速搜索和注释序列的算法。生物信息学之外的一个应用例子是在自然语言处理和程序编译中使用多上下文无关语法进行解析。在整个项目中,学生都将积极参与。当实现完成后,学生将帮助PI呈现这些实现,并将其用于项目网页。本项目涉及基础计算机科学理论与生物信息学的应用。它将在自动机和形式语言方面产生新的知识和案例研究结果,这是计算机科学课程的重要组成部分。它也将帮助参与的学生掌握这些主题。本项目将通过以下方法开发蛋白质和RNA序列的搜索、比较和注释算法:(1)基于种子的匹配:为了找到与给定序列的近似匹配,流行的比对和搜索工具BLAST首先定位固定长度区域(种子)的精确匹配,并在种子周围扩展匹配区域。该项目以新颖的方式推广了种子的使用,以解决正则表达式和上下文无关语法描述模式的模式匹配和注释问题。初步结果表明,所提出的基于种子的方法在2MB文本上找到匹配的速度比UNIX GREP实用程序快2.5倍;(2)后缀树/数组匹配:对于固定字母上有限长模式的标注问题,本项目建议使用后缀树(或后缀数组)扩展附加信息,以便从候选集、正则表达式或上下文无关的语法中识别生成给定字符串的语法;(3) RNA的新表示:本项目提出了一种新的RNA二级结构表示,其中二维结构信息嵌入到具有所需特征的序列中。对于RNA序列,提出的新算法将利用这种表示的优势进行快速RNA搜索,注释和比较(多个RNA定位共同子结构)。第一年将解决基于种子的搜索和RNA比较问题(使用新的表示),第二年将解决注释问题,因为序列注释将使用第一年的结果。该项目将从PROSITE、Rfam、RNA STRAND、rCAD(特别是RNA序列)等公开可用的数据库中创建实验数据库。Bpseq文件),用于解释和显示所开发算法的结果。每学期,在项目期间,学生将参与开发、实施、测试新的算法、用户界面和相关的支持工具。
英文摘要
This project aims to develop fast algorithms for searching, comparing, and annotating protein and RNA sequences. Regular expression matching is commonly used in UNIX-based systems for searching texts, and in PROSITE website for searching patterns in protein sequences. Context free grammar matching is used in searching RNA sequences. Annotating biological sequences (DNA, protein, and RNA sequences) using regular expression and context free grammar described motifs is an important application available in many public databases, websites, and software tools (e.g. PROSITE, Locomotif). The results of the proposed project will be helpful for programmers who develop pattern matching and parsing applications which would benefit from fast algorithms for searching and annotating sequences. An example of such an application outside bioinformatics is parsing with multiple context free grammars, which is used in natural language processing and program compiling. There will be strong student involvement during the entire project. As implementations become complete, students will help the PI present these implementations, and make them available for use in project web pages. This project involves fundamental computer science theory with applications in bioinformatics. It will yield new knowledge and case study results in automata and formal languages which are essential parts of the computer science curriculum. It will also help involved students master these topics.This project will develop algorithms for searching, comparing and annotating protein and RNA sequences by using (1) Seed-based matching: For finding an approximate match to a given sequence, popular alignment and search tool BLAST locates first an exact match of fixed length region (seed) and extends the matching region around the seed. This project generalizes the use of seeds in novel ways to pattern matching and annotation problems for regular expression and context free grammar described patterns. Initial results indicate that the proposed seed-based approach finds matches about 2.5 times faster than UNIX GREP utility on 2MB texts; (2) Suffix tree/array-based matching: For the annotation problem with bounded-length patterns over fixed alphabets, this project proposes using a suffix tree (or a suffix array) extended with additional information in order to identify from a candidate set, a regular expression or a context free grammar that generates a given string; and (3) A new representation for RNA: This project proposes a new RNA secondary structure representation in which two-dimensional structure information is embedded in the sequence with desirable features. For RNA sequences, the proposed new algorithms will exploit the advantages of this representation for fast RNA search, annotation, and comparison (of multiple RNAs to locate common substructures). Seed-based search and RNA comparison problems (using the new representation) will be addressed in the first year, and annotation problems in the second year as sequence annotation will make use of the results from the first year. The project will create experimental databases from publicly available databases such as PROSITE, Rfam, RNA STRAND, rCAD (in particular RNA sequences in .bpseq files) for the purpose of explaining and showing the results of the developed algorithms. Every semester, during the project, students will be involved in developing, implementing, testing new algorithms, user interfaces, and relevant support tools.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: