Selecting signature oligonucleotides to identify organisms using DNA arrays

Selecting signature oligonucleotides to identify organisms using DNA arrays
复制标题

DOI:
10.1093/bioinformatics/18.10.1340
复制
发表时间:
2002-10-01
期刊:
影响因子:
5.8
通讯作者:
Schliep, A
Schliep, A
中科院分区:
生物学3区
文献类型:
--
作者:
Kaderali, L;Schliep, A

文献摘要

被引文献

相似文献

动机:DNA 阵列是一种非常有用的工具,可以快速识别某些给定样本中存在的生物制剂,例如样本。识别引起疾病的病毒,用于食品工业的质量控制,或确定污染饮用水的细菌。选择附着在阵列表面的特定寡核苷酸是实验设计过程中的一个相关问题。给定一组 S 组基因组序列(目标序列),任务是为 S 中的每个序列找到至少一个寡核苷酸(称为探针)。该探针将附着在阵列表面,并且必须以不会与除预期目标序列之外的任何其他序列杂交的方式进行选择。此外,阵列上的所有探针必须在相同的反应条件下与其预期目标杂交,最重要的是在进行实验的温度 T 下。结果:我们针对探针设计问题提出了一种有效的算法。使用扩展的最近邻模型计算所有可能的探针-靶标相互作用的解链温度,允许双链体中的非沃森-克里克碱基配对和未配对碱基。为了有效地计算温度,引入了后缀树和基于动态规划的对齐算法的组合。预处理期间的额外过滤步骤提高了计算速度。两个案例研究证明了算法的实用性:HIV-1 亚型的识别以及来自大于或等于 400 个生物体的 28S rDNA 序列的识别。
Motivation: DNA arrays are a very useful tool to quickly identify biological agents present in some given sample, e.g. to identify viruses causing disease, for quality control in the food industry, or to determine bacteria contaminating drinking water. The selection of specific oligos to attach to the array surface is a relevant problem in the experiment design process. Given a set S of genomic sequences (the target sequences), the task is to find at least one oligonucleotide, called probe, for each sequence in S. This probe will be attached to the array surface, and must be chosen in a way that it will not hybridize to any other sequence but the intended target. Furthermore, all probes on the array must hybridize to their intended targets under the same reaction conditions, most importantly at the temperature T at which the experiment is conducted.Results: We present an efficient algorithm for the probe design problem. Melting temperatures are calculated for all possible probe-target interactions using an extended nearest-neighbor model, allowing for both non-Watson-Crick base-pairing and unpaired bases within a duplex. To compute temperatures efficiently, a combination of suffix trees and dynamic programming based alignment algorithms is introduced. Additional filtering steps during preprocessing increase the speed of the computation.The practicability of the algorithms is demonstrated by two case studies: The identification of HIV-1 subtypes, and of 28S rDNA sequences from greater than or equal to400 organisms.