Data mining and visualisation: general discussion.
Data mining and visualisation: general discussion.
复制标题
数据挖掘和可视化:一般讨论。
DOI:
10.1039/c9fd90044f
复制
发表时间:
2019
影响因子:
3.4
通讯作者:
Afonso C
中科院分区:
文献类型:
--
作者:
Afonso C
Gianluca Tri ro opened discussion of the paper by Johan Trygg: With regard to the multivariate analysis of data with different sizes, structures and sources, could the kind of transformation on the raw data affect the nal distribution, possibly misleading the interpretation of the correlation matrix? Also, how much do the outliers weigh on the entire multiblock analysis?Johan Trygg answered: Of course, non-linear data transformations will have an in uence on the correlation matrix and hence the resulting multivariate models. The reason why transformation or any preprocessing is made is to correct for “unwanted” shape/variability or to linearize the data. Typically today, data processing work ows include multiple steps and it is hard to understand the impact or interaction between those transformation steps, and how it affects the outcome. Normally, this is not a problem as the work ows are quite established, but otherwise we recommend using a multiblock analysis technique like Joint and Unique MultiBlock Analysis (JUMBA) 1 where the data set in each step in the work ow is assembled and modeled together. Think of principal component analysis (PCA), but on multiple blocks of data, where you get an overview, and can visualize the ow of data and how it changes in each step in a single visualization plot. Multiblock models can also be used for comparing the effect between transformed and nontransformed data. In addition, it is quite common to analyze the same sample on multiple instruments, hence providing multi-block data where JUMBA is great.