Data mining and visualisation: general discussion.

Data mining and visualisation: general discussion.
复制标题

数据挖掘和可视化:一般讨论。

DOI:
10.1039/c9fd90044f
复制
发表时间:
2019
影响因子:
3.4
通讯作者:
Afonso C
Afonso C
中科院分区:
化学2区
文献类型:
--
作者:
Afonso C

文献摘要

被引文献

相似文献

GianLuca Triro开始讨论Johan Trygg的论文:对于具有不同大小、结构和来源的数据的多变量分析,原始数据上的那种转换是否会影响NAL分布,从而可能误导相关矩阵的解释?另外,离群值在整个多块分析中的权重有多大?Johan Trygg回答:当然,非线性数据转换会对相关矩阵产生影响,从而产生多变量模型。之所以要进行变换或任何预处理,是为了纠正“不需要的”形状/可变性,或者使数据线性化。通常今天,数据处理工作OW包括多个步骤,很难理解这些转换步骤之间的影响或相互作用,以及它如何影响结果。通常情况下,这不是问题,因为WorkOW是相当成熟的,但否则我们建议使用多块分析技术,如联合和唯一多块分析(JUMBA)1,其中Workow中每个步骤中的数据集都被组合在一起并进行建模。想想主成分分析(),但是在多个数据块上,你可以得到一个概述,可以可视化数据的主成分分析,以及它在单个可视化绘图中每一步是如何变化的。多区块模型也可用于比较变换数据和未变换数据之间的效果。此外,在多个仪器上分析同一样本是很常见的,因此在JUMBA很好的地方提供了多块数据。
Gianluca Tri ro opened discussion of the paper by Johan Trygg: With regard to the multivariate analysis of data with different sizes, structures and sources, could the kind of transformation on the raw data affect the nal distribution, possibly misleading the interpretation of the correlation matrix? Also, how much do the outliers weigh on the entire multiblock analysis?Johan Trygg answered: Of course, non-linear data transformations will have an in uence on the correlation matrix and hence the resulting multivariate models. The reason why transformation or any preprocessing is made is to correct for “unwanted” shape/variability or to linearize the data. Typically today, data processing work ows include multiple steps and it is hard to understand the impact or interaction between those transformation steps, and how it affects the outcome. Normally, this is not a problem as the work ows are quite established, but otherwise we recommend using a multiblock analysis technique like Joint and Unique MultiBlock Analysis (JUMBA) 1 where the data set in each step in the work ow is assembled and modeled together. Think of principal component analysis (PCA), but on multiple blocks of data, where you get an overview, and can visualize the ow of data and how it changes in each step in a single visualization plot. Multiblock models can also be used for comparing the effect between transformed and nontransformed data. In addition, it is quite common to analyze the same sample on multiple instruments, hence providing multi-block data where JUMBA is great.