Splice site identification by idlBNs

Splice site identification by idlBNs
复制标题

DOI:
10.1093/bioinformatics/bth932
复制
发表时间:
2004-08-04
期刊:
影响因子:
5.8
通讯作者:
Guigo, Roderic
Guigo, Roderic
中科院分区:
生物学3区
文献类型:
--
作者:
Castelo, Robert;Guigo, Roderic

文献摘要

被引文献

相似文献

动机:核苷酸序列中功能位点的计算识别是用于分析基因组数据的许多算法的核心。该识别基于从训练集估计的统计参数。通常,由于参数数量庞大,很难获得一致的估计量。为了简化估计问题,在沿着位点的核苷酸之间施加独立的假设。然而,这可能会限制估计误差的最小值。结果:在本文中,我们介绍了一种新的方法,在识别功能位点的背景下,发现一组合理的独立性假设支持的数据,核苷酸之间,并使用它来执行识别的网站由他们的似然比。更重要的是,在许多实际情况下,随着训练样本量的增加,它能够提高其性能。我们将该方法应用于剪接位点的识别,并进一步评估其在外显子和基因预测方面的效果。
Motivation: Computational identification of functional sites in nucleotide sequences is at the core of many algorithms for the analysis of genomic data. This identification is based on the statistical parameters estimated from a training set. Often, because of the huge number of parameters, it is difficult to obtain consistent estimators. To simplify the estimation problem, one imposes independent assumptions between the nucleotides along the site. However, this can potentially limit the minimum value of the estimation error.Results: In this paper, we introduce a novel method in the context of identifying functional sites, that finds a reasonable set of independence assumptions supported by the data, among the nucleotides, and uses it to perform the identification of the sites by their likelihood ratio. More importantly, in many practical situations it is capable of improving its performance as the training sample size increases. We apply the method to the identification of splice sites, and further evaluate its effect within the context of exon and gene prediction.