Data Fusion in Metabolomics Using Coupled Matrix and Tensor Factorizations

Data Fusion in Metabolomics Using Coupled Matrix and Tensor Factorizations
复制标题

DOI:
10.1109/jproc.2015.2438719
复制
发表时间:
2015-09-01
影响因子:
20.6
通讯作者:
Smilde, Age K.
Smilde, Age K.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Acar, Evrim;Bro, Rasmus;Smilde, Age K.

文献摘要

被引文献

相似文献

代谢组学的目标是确定与某些条件或疾病相关的生物标记物/模式,重点是使用核磁共振(核磁共振)光谱、液质联用(LC-MS)和荧光光谱等多种分析技术检测生物样品中的化学物质,如尿液和血液。使用这些方法测量的数据集提供了部分补充信息,它们的联合分析有可能揭示潜在的结构,否则很难提取。虽然我们可以使用不同的分析方法收集大量数据,但数据融合仍然是一项具有挑战性的任务,特别是当目标是捕捉潜在因素并将其用于解释时,例如用于生物标志物识别。此外,许多数据融合应用需要对具有共享/非共享因子的异类(即,以高阶张量和矩阵的形式)数据集进行联合分析。为了联合分析这类异质数据集,我们将数据融合问题描述为一个耦合的矩阵和张量分解(CMTF)问题,并讨论了它对揭示结构的数据融合模型的扩展,即能够识别共享和非共享因素的数据融合模型。在存在共享/非共享因子的情况下,传统的数据融合方法通常是基于矩阵分解的方法。使用模拟和典型的实验耦合数据集,我们评估了各种最新的数据融合方法的性能,并证明了尽管基于矩阵因式分解的方法在用于异质数据集的联合分析时存在局限性,但揭示结构的CMTF模型可以通过利用高阶数据集的低阶结构来成功地捕捉潜在因素。
With a goal of identifying biomarkers/patterns related to certain conditions or diseases, metabolomics focuses on the detection of chemical substances in biological samples such as urine and blood using a number of analytical techniques, including nuclear magnetic resonance (NMR) spectroscopy, liquid chromatography-mass spectrometry (LC-MS), and fluorescence spectroscopy. Data sets measured using these methods provide partly complementary information, and their joint analysis has the potential to reveal underlying structures, which are, otherwise, difficult to extract. While we can collect vast amounts of data using different analytical methods, data fusion remains a challenging task, in particular, when the goal is to capture the underlying factors and use them for interpretation, e.g., for biomarker identification. Furthermore, many data fusion applications require joint analysis of heterogeneous (i.e., in the form of higher order tensors and matrices) data sets with shared/unshared factors. In order to jointly analyze such heterogeneous data sets, we formulate data fusion as a coupled matrix and tensor factorization (CMTF) problem, which has already proved useful in many data mining applications, and discuss its extension to a structure-revealing data fusion model, i.e., a data fusion model that can identify shared and unshared factors. The traditional methods commonly used for data fusion in the presence of shared/unshared factors are matrix factorization-based methods. Using both simulations and prototypical experimental coupled data sets, we assess the performance of various state-of-the-art data fusion methods and demonstrate that while matrix factorization-based approaches have limitations when used for joint analysis of heterogeneous data sets, the structure-revealing CMTF model can successfully capture the underlying factors by exploiting the low-rank structure of higher order data sets.