Kangaroo--a pattern-matching program for biological sequences.

Kangaroo--a pattern-matching program for biological sequences.
复制标题

DOI:
10.1186/1471-2105-3-20
复制
发表时间:
2002-07-31
期刊:
影响因子:
3
通讯作者:
Hogue CW
Hogue CW
中科院分区:
生物学4区
文献类型:
--
作者:
Betel D;Hogue CW

文献摘要

参考文献

被引文献

相似文献

生物学家通常对执行简单的数据库搜索以识别包含明确定义的序列模式的蛋白质或基因感兴趣。许多数据库不提供直接或容易获得的查询工具来执行简单的搜索,例如识别转录结合位点、蛋白质基序或重复DNA序列。然而,在许多情况下,简单的模式匹配搜索可以揭示大量信息。在本文中,我们提出了一种正则表达式模式匹配工具,用于识别人类编码区的短重复DNA序列,以识别错配修复缺陷细胞中的潜在突变位点。Kangaroo是一个基于Web的正则表达式模式匹配程序,可以搜索十种不同生物体中的DNA、蛋白质或编码区序列的模式。该计划的实施,以促进广泛的查询,没有限制的长度或复杂性的查询表达式。该程序可在http://bioinfo.mshri.on.ca/kangaroo/上访问,源代码可在http://sourceforge.net/projects/slritools/上免费分发。一个低级的简单模式匹配应用程序可以证明是许多研究环境中的有用工具。例如,Kangaroo被用于鉴定人结肠直肠癌变体中的潜在遗传靶标,该变体的特征在于含有单核苷酸重复序列的编码区中的高频率突变。
Biologists are often interested in performing a simple database search to identify proteins or genes that contain a well-defined sequence pattern. Many databases do not provide straightforward or readily available query tools to perform simple searches, such as identifying transcription binding sites, protein motifs, or repetitive DNA sequences. However, in many cases simple pattern-matching searches can reveal a wealth of information. We present in this paper a regular expression pattern-matching tool that was used to identify short repetitive DNA sequences in human coding regions for the purpose of identifying potential mutation sites in mismatch repair deficient cells. Kangaroo is a web-based regular expression pattern-matching program that can search for patterns in DNA, protein, or coding region sequences in ten different organisms. The program is implemented to facilitate a wide range of queries with no restriction on the length or complexity of the query expression. The program is accessible on the web at http://bioinfo.mshri.on.ca/kangaroo/ and the source code is freely distributed at http://sourceforge.net/projects/slritools/. A low-level simple pattern-matching application can prove to be a useful tool in many research settings. For example, Kangaroo was used to identify potential genetic targets in a human colorectal cancer variant that is characterized by a high frequency of mutations in coding regions containing mononucleotide repeats.
DOI: 10.1093/nar/20.11.2871
发表时间: 1992-06-11
影响因子: 14.9
作者:
PESOLE, G;PRUNELLA, N;SACCONE, C
通讯作者: SACCONE, C
DOI: 10.1016/0097-8485(93)85006-x
发表时间: 1993-06-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
WOOTTON, JC;FEDERHEN, S
通讯作者: FEDERHEN, S
DOI: 10.1093/embo-reports/kvd092
发表时间: 2000-11-01
期刊: EMBO REPORTS
影响因子: 7.7
作者:
Cokol, M;Nair, R;Rost, B
通讯作者: Rost, B
DOI: 10.1016/s0168-9525(97)01347-4
发表时间: 1997-12-01
期刊: TRENDS IN GENETICS
影响因子: 11.4
作者:
Dsouza, M;Larsen, N;Overbeek, R
通讯作者: Overbeek, R
DOI: 10.1093/bioinformatics/16.5.439
发表时间: 2000-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pesole, G;Liuni, S;D'Souza, M
通讯作者: D'Souza, M