Improving learning accuracy by using synthetic samples for small datasets with non-linear attribute dependency

Improving learning accuracy by using synthetic samples for small datasets with non-linear attribute dependency
复制标题

DOI:
10.1016/j.dss.2013.12.007
复制
发表时间:
2014-03
期刊:
Decis. Support Syst.
影响因子:
--
通讯作者:
Der-Chiang Li;Liang-Sian Lin;Li-Jhong Peng
Der-Chiang Li;Liang-Sian Lin;Li-Jhong Peng
中科院分区:
其他
文献类型:
--
作者:
Der-Chiang Li;Liang-Sian Lin;Li-Jhong Peng

文献摘要

被引文献

相似文献

小数据问题通常在新制造过程的早期阶段遇到,这给学者和从业者带来了挑战,因为当缺乏足够的数据时,学习模型很难实现良好的性能。虚拟样本生成(VSG)已被证明是一种有效的方法来克服这个问题在各种领域的广泛的研究。这些工作通常假设属性之间的关系是相互独立的,并通过使用这些样本分布产生合成数据。然而,如果真实的数据具有相关属性,则VSG技术可能是无效的。因此,本研究提供了一种新的方法来产生相关的虚拟样本与非线性属性依赖。为了建立独立属性和依赖属性之间的关系模型,我们采用基因表达式编程(GEP)来寻找最合适的数学模型。通过一个实际数据集和三个真实的UCI数据集验证了该方法的有效性,结果表明,该方法对反向传播神经网络(BPN)的学习精度优于传统的大趋势扩散(MTD)和多元回归分析(MRA)方法.
Small-data problems are commonly encountered in the early stages of a new manufacturing procedure, presenting challenges to both academics and practitioners, as good performance is difficult to achieve with learning models when there is a lack of sufficient data. Virtual sample generation (VSG) has been shown to be an effective method to overcome this issue in a wide range of studies in various fields. Such works usually assume that the relations among attributes are independent of each other, and produce synthetic data by using sample distributions of these. However, the VSG technique may be ineffective if the real data has interrelated attributes. Therefore, this research provides a novel procedure to generate related virtual samples with non-linear attribute dependency. To construct a relational model between the independent and dependent attributes, we employ gene expression programming (GEP) to find the most suitable mathematical model. One practical dataset and three real UCI datasets are presented in this paper to verify the effectiveness of the proposed method, and the results show that the proposed approach has better learning accuracy with regard to a back-propagation neural (BPN) network than that of the well-known mega-trend-diffusion (MTD) and the multi regression analysis (MRA) approaches.