Reducing the effect of sample bias for small data sets with double-weighted support vector transfer regression

Reducing the effect of sample bias for small data sets with double-weighted support vector transfer regression
复制标题

DOI:
10.1111/mice.12617
复制
发表时间:
2020-09-01
影响因子:
9.6
通讯作者:
Paal, Stephanie German
Paal, Stephanie German
中科院分区:
工程技术1区
文献类型:
--
作者:
Luo, Huan;Paal, Stephanie German

文献摘要

被引文献

相似文献

小数据集在机器学习(ML)领域是一个极具挑战性的问题,具体来说,在回归场景中,由于缺乏相关数据可能导致ML模型具有很大的偏差。然而,有许多应用程序,其中一个纯粹的数据驱动的过程将是有利的,但大量的数据不可用。本文提出了一种新的基于回归的迁移学习(TL)模型来解决这一挑战,其中TL被定义为从大型相关数据集(源域数据)到小型数据集(目标域数据)的知识迁移。提出的TL模型被称为双加权支持向量转移回归(DW-SVTR),它耦合最小二乘支持向量机回归(LS-SVMR)与两个权重函数。第一权重函数使用核均值匹配(KMM)来重新加权源域数据,使得源域数据和目标域数据在再现核希尔伯特空间(RKHS)中的均值接近。以这种方式,与目标域点相关的源域数据点具有比不相关的源域点更大的权重。第二个权值是估计残差的函数,其目的是进一步减少不相关源域点的负干扰。所提出的方法进行了评估和验证,通过模拟数据和增强的抗剪强度预测的基础上有限的非延性列数据。具体而言,后者的结果表明,建议DW-SVTR可以减少均方根误差(RMSE)的34%,提高决定系数(R-2)的229%。这些数值结果表明,DW-SVTR显着降低了小样本偏差的影响,并提高了预测性能相比,标准ML方法。
Small data sets are an extremely challenging problem in the machine learning (ML) realm, and in specific, in regression scenarios, as the lack of relevant data can lead to ML models that have large bias. However, there are many applications for which a purely data-driven procedure would be advantageous, but a large amount of data are not available. This article proposes a novel regression-based transfer learning (TL) model to address this challenge, where TL is defined as knowledge transfer from a large, relevant data set (source domain data) to a small data set (target domain data). The proposed TL model is termed double-weighted support vector transfer regression (DW-SVTR), which couples least squares support vector machines for regression (LS-SVMR) with two weight functions. The first weight function uses kernel mean matching (KMM) to reweight the source domain data such that the mean values of the source and target domain data in a reproduced kernel Hilbert space (RKHS) are close. In this way, the source domain data points relevant to the target domain points have a larger weight than irrelevant source domain points. The second weight is a function of estimated residuals, which aims to further reduce the negative interference of irrelevant source domain points. The proposed approach is assessed and validated via simulated data and by enhanced shear strength prediction of nonductile columns based on limited availability of nonductile column data. Specifically, the results for the latter show that the proposed DW-SVTR can reduce the root mean square error (RMSE) by 34% and enhance the coefficient of determination (R-2) by 229%. These numerical results demonstrate that the DW-SVTR significantly reduces the effect of small sample bias and improves prediction performance compared to standard ML methods.