Focused Stochastic Neighbor Embedding for Better Preserving Points of Interest

Focused Stochastic Neighbor Embedding for Better Preserving Points of Interest
复制标题

DOI:
10.1109/bdcat56447.2022.00043
复制
发表时间:
2022-12
期刊:
2022 IEEE/ACM International Conference on Big Data Computing, Applications and Technologies (BDCAT)
影响因子:
--
通讯作者:
Rafael Baez Ramirez;Sanuj Kumar;Tuan M. V. Le;H. Cao
Rafael Baez Ramirez;Sanuj Kumar;Tuan M. V. Le;H. Cao
中科院分区:
其他
文献类型:
--
作者:
Rafael Baez Ramirez;Sanuj Kumar;Tuan M. V. Le;H. Cao

文献摘要

相似文献

抽象性约简的目的是找到高维数据的低维嵌入,使得低维表示保留原始数据中结构的一些有意义的属性。当低维空间是2维或3维时,可以使用散点图来可视化低维嵌入。大多数现有的方法试图保持所有数据点的局部邻域。然而,一般来说,不可能为低维空间中的所有数据点保留所有这些信息。因此,可能会有一些数据点,其邻域由于信息丢失而没有如实地显示在可视化中。如果信息丢失发生在一组特定的兴趣点周围(例如,特定患者或观察到的蛋白质),这可能是有问题的,因为撤回的见解对于这些观察到的数据点可能不准确。因此,在本文中,我们引入了一个称为集中降维的问题,在给定原始高维数据集和一组兴趣点的情况下,我们希望找到原始数据的2维或3维嵌入,以使局部信息损失尽可能最小化兴趣点周围的社区。换句话说,如果信息丢失是不可避免的,它不应该发生在兴趣点周围。为了解决这个问题,我们扩展了随机邻居嵌入方法,并引入了一个集中的目标函数,我们把更多的权重放在涉及兴趣点的损失上。在真实数据集上的实验表明,该方法能更好地保持兴趣点的局部邻域结构,同时生成的可视化效果与随机邻域嵌入方法一样好。
Dimensionality reduction aims to find low-dimensional embeddings of high-dimensional data such that the low-dimensional representation preserves some meaningful properties of structures in the original data. When low-dimensional space is 2- or 3-dimensional, the low-dimensional embeddings can be visualized using a scatterplot map. Most of the existing methods try to preserve the local neighborhoods of all data points. However, in general, it is impossible to retain all such information for all data points in the low-dimensional space. As a result, there could be some data points whose neighborhoods are not faithfully displayed in the visualization due to information loss. If the information loss happens around a specific set of points of interest (e.g., specific patients, or proteins under observed), this may be problematic because the withdrawn insights may not be accurate for these observed data points. Therefore, in this paper, we introduce a problem called focused dimensionality reduction where given an original high-dimensional dataset and a set of points of interest, we want to find 2- or 3-dimensional embeddings of the original data such that the information loss in the local neighborhoods surrounding the points of interest is minimized as much as possible. In other words, if the information loss is inevitable, it should not happen around the points of interest. To solve the problem, we extend the stochastic neighbor embedding method and introduce a focused objective function where we put more weight on losses that involve points of interest. Experiments on real-world datasets show that our proposed method is better in preserving the local neighborhood structure of points of interest while the generated visualizations are as good as those generated by the stochastic neighbor embedding method.