Compression-based classification of biological sequences and structures via the Universal Similarity Metric: experimental assessment.

Compression-based classification of biological sequences and structures via the Universal Similarity Metric: experimental assessment.
复制标题

DOI:
10.1186/1471-2105-8-252
复制
发表时间:
2007-07-13
期刊:
影响因子:
3
通讯作者:
Valiente G
Valiente G
中科院分区:
生物学4区
文献类型:
--
作者:
Ferragina P;Giancarlo R;Greco V;Manzini G;Valiente G

文献摘要

参考文献

被引文献

相似文献

序列相似性是生物学分类和系统发育研究中的一个重要数学概念。目前主要使用对齐处理。然而,比对方法似乎不适合后基因组研究,因为它们不能很好地与数据集大小相匹配,并且它们似乎仅限于基因组和蛋白质组序列。因此,积极追求无重复的相似性度量。其中,USM(Universal Similarity Metric,通用相似性度量)获得了突出地位。它以柯尔莫哥洛夫复杂性理论为基础,普适性是其最新颖的显著特征。由于它只能通过数据压缩近似,USM是一种方法,而不是量化两个字符串相似性的公式。USM有三种近似,即UCD(通用压缩差异),NCD(归一化压缩差异)和CD(压缩差异)。他们的适用性和鲁棒性测试各种数据集产生的第一个大规模的定量估计,USM方法及其近似值的价值。尽管围绕USM开发了丰富的理论,但其实验评估具有局限性:只有少数数据压缩器与USM一起进行了测试,并且大多数是在定性水平上,UCD,NCD和CD之间没有比较,并且USM与现有方法之间没有比较,无论是基于对齐还是不基于对齐,似乎都是可用的。我们实验测试的USM方法,使用25个压缩机,所有三个已知的近似值和六个数据集的相关分子生物学。这提供了这种方法的第一个系统和定量的实验评估,自然补充了许多理论和初步的实验结果。此外,我们比较了USM方法与基于对齐的方法。我们可以把实验分成两组。第一个,通过ROC(受试者工作曲线)分析,旨在评估的方法来区分和分类生物序列和结构的内在能力。第二组实验旨在评估两种常用的分类算法UPGMA(具有算术平均值的未加权对组方法)和NJ(邻居连接)可以如何使用该方法来执行它们的任务,它们的性能针对黄金标准并使用众所周知的统计指标来评估,即,F测度和划分距离。基于实验,可以得出几个结论,并从他们,新的有价值的指导方针使用USM的生物数据。主要的报告如下。UCD和NCD是不可区分的,即,在交叉实验和数据集上,它们产生的统计指标值几乎相同,而CD几乎总是比两者更差。UPGMA似乎相对于NJ产生更好的分类结果,即,在大量实验、压缩机和USM近似选择中,统计指标的值更好(差异10%或以上)。基于PPM(部分匹配预测)的压缩程序PPMd(用于通用数据)和Gencompress(用于DNA)是我们使用的压缩算法中表现最好的,尽管通过统计指标测量的性能差异,它们与其他算法之间的差异主要取决于数据集,可能没有预期的那么大。与UCD或NCD和UPGMA一起使用的PPMd在序列数据上非常接近,尽管在比对方法的性能上更差(在F测量上的差异小于2%)。然而,它可以很好地扩展数据集的大小,并且可以处理序列以外的数据。总之,我们的定量分析自然补充了USM背后的丰富理论,并支持该方法值得使用的结论,因为它具有鲁棒性,灵活性,可扩展性和与现有技术的竞争力。特别是,该方法适用于文本格式的所有生物数据。软件和数据集可在补充材料网页上根据GNU GPL获得。
Similarity of sequences is a key mathematical notion for Classification and Phylogenetic studies in Biology. It is currently primarily handled using alignments. However, the alignment methods seem inadequate for post-genomic studies since they do not scale well with data set size and they seem to be confined only to genomic and proteomic sequences. Therefore, alignment-free similarity measures are actively pursued. Among those, USM (Universal Similarity Metric) has gained prominence. It is based on the deep theory of Kolmogorov Complexity and universality is its most novel striking feature. Since it can only be approximated via data compression, USM is a methodology rather than a formula quantifying the similarity of two strings. Three approximations of USM are available, namely UCD (Universal Compression Dissimilarity), NCD (Normalized Compression Dissimilarity) and CD (Compression Dissimilarity). Their applicability and robustness is tested on various data sets yielding a first massive quantitative estimate that the USM methodology and its approximations are of value. Despite the rich theory developed around USM, its experimental assessment has limitations: only a few data compressors have been tested in conjunction with USM and mostly at a qualitative level, no comparison among UCD, NCD and CD is available and no comparison of USM with existing methods, both based on alignments and not, seems to be available. We experimentally test the USM methodology by using 25 compressors, all three of its known approximations and six data sets of relevance to Molecular Biology. This offers the first systematic and quantitative experimental assessment of this methodology, that naturally complements the many theoretical and the preliminary experimental results available. Moreover, we compare the USM methodology both with methods based on alignments and not. We may group our experiments into two sets. The first one, performed via ROC (Receiver Operating Curve) analysis, aims at assessing the intrinsic ability of the methodology to discriminate and classify biological sequences and structures. A second set of experiments aims at assessing how well two commonly available classification algorithms, UPGMA (Unweighted Pair Group Method with Arithmetic Mean) and NJ (Neighbor Joining), can use the methodology to perform their task, their performance being evaluated against gold standards and with the use of well known statistical indexes, i.e., the F-measure and the partition distance. Based on the experiments, several conclusions can be drawn and, from them, novel valuable guidelines for the use of USM on biological data. The main ones are reported next. UCD and NCD are indistinguishable, i.e., they yield nearly the same values of the statistical indexes we have used, accross experiments and data sets, while CD is almost always worse than both. UPGMA seems to yield better classification results with respect to NJ, i.e., better values of the statistical indexes (10% difference or above), on a substantial fraction of experiments, compressors and USM approximation choices. The compression program PPMd, based on PPM (Prediction by Partial Matching), for generic data and Gencompress for DNA, are the best performers among the compression algorithms we have used, although the difference in performance, as measured by statistical indexes, between them and the other algorithms depends critically on the data set and may not be as large as expected. PPMd used with UCD or NCD and UPGMA, on sequence data is very close, although worse, in performance with the alignment methods (less than 2% difference on the F-measure). Yet, it scales well with data set size and it can work on data other than sequences. In summary, our quantitative analysis naturally complements the rich theory behind USM and supports the conclusion that the methodology is worth using because of its robustness, flexibility, scalability, and competitiveness with existing techniques. In particular, the methodology applies to all biological data in textual format. The software and data sets are available under the GNU GPL at the supplementary material web page.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1186/1748-7188-1-4
发表时间: 2006-01-01
影响因子: 1
作者:
Apostolico, Alberto;Comin, Matteo;Parida, Laxmi
通讯作者: Parida, Laxmi
DOI: 10.1145/950620.950622
发表时间: 2003-11-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
Buchsbaum, AL;Fowler, GS;Giancarlo, R
通讯作者: Giancarlo, R
DOI: 10.1109/51.940049
发表时间: 2001-07-01
影响因子: --
作者:
Chen, X;Kwong, S;Li, M
通讯作者: Li, M
DOI: 10.1109/tit.2004.838101
发表时间: 2004-12-01
影响因子: 2.5
作者:
Li, M;Chen, X;Vitányi, PMB
通讯作者: Vitányi, PMB