A kernel-based approach for detecting outliers of high-dimensional biological data.

A kernel-based approach for detecting outliers of high-dimensional biological data.
复制标题

一种基于内核的方法,用于检测高维生物数据的异常值。

DOI:
10.1186/1471-2105-10-s4-s7
复制
发表时间:
2009-04-29
期刊:
影响因子:
3
通讯作者:
Gao J
Gao J
中科院分区:
生物学4区
文献类型:
--
作者:
Oh JH;Gao J

文献摘要

被引文献

相似文献

在许多情况下,生物医学数据集包含异常值,导致难以实现可靠的知识发现。不去除异常值的数据分析可能会导致错误的结果并提供误导性信息。我们提出了一种基于 Kullback-Leibler (KL) 散度的新异常值检测方法。 KL 散度的最初概念被设计为两个分布之间距离的度量。由此,我们通过形成由最近邻居组成的样本集将其扩展到生物样本异常值检测。 KL 散度是在有和没有测试样本的两个样本集之间定义的。为了处理样本分布的非线性,原始数据被映射到更高的特征空间。我们解决了 KL 散度计算过程中由于样本量较小而导致的奇异性问题。应用内核函数以避免直接使用映射函数。该方法的性能在一个合成数据集、两个公共微阵列数据集和一个用于肝癌研究的质谱数据集上得到了证明。与基于马哈拉诺比斯距离的方法和一类支持向量机(SVM)的比较研究表明,所提出的方法在发现异常值方面表现更好。我们的想法源自马尔可夫毯子算法,这是一种基于 KL 散度的特征选择方法。也就是说,虽然马尔可夫毯子算法删除了冗余和不相关的特征,但我们提出的方法检测异常值。与其他算法相比,我们提出的方法对于小样本和高维生物数据表现出更好或相当的性能。这表明所提出的方法可用于检测生物数据集中的异常值。
In many cases biomedical data sets contain outliers that make it difficult to achieve reliable knowledge discovery. Data analysis without removing outliers could lead to wrong results and provide misleading information. We propose a new outlier detection method based on Kullback-Leibler (KL) divergence. The original concept of KL divergence was designed as a measure of distance between two distributions. Stemming from that, we extend it to biological sample outlier detection by forming sample sets composed of nearest neighbors. KL divergence is defined between two sample sets with and without the test sample. To handle the non-linearity of sample distribution, original data is mapped into a higher feature space. We address the singularity problem due to small sample size during KL divergence calculation. Kernel functions are applied to avoid direct use of mapping functions. The performance of the proposed method is demonstrated on a synthetic data set, two public microarray data sets, and a mass spectrometry data set for liver cancer study. Comparative studies with Mahalanobis distance based method and one-class support vector machine (SVM) are reported showing that the proposed method performs better in finding outliers. Our idea was derived from Markov blanket algorithm that is a feature selection method based on KL divergence. That is, while Markov blanket algorithm removes redundant and irrelevant features, our proposed method detects outliers. Compared to other algorithms, our proposed method shows better or comparable performance for small sample and high-dimensional biological data. This indicates that the proposed method can be used to detect outliers in biological data sets.