Deep learning with small datasets: using autoencoders to address limited datasets in construction management

Deep learning with small datasets: using autoencoders to address limited datasets in construction management
复制标题

DOI:
10.1016/j.asoc.2021.107836
复制
发表时间:
2021-11
期刊:
Appl. Soft Comput.
影响因子:
--
通讯作者:
Juan Manuel Davila Delgado;Lukumon O. Oyedele
Juan Manuel Davila Delgado;Lukumon O. Oyedele
中科院分区:
其他
文献类型:
--
作者:
Juan Manuel Davila Delgado;Lukumon O. Oyedele

文献摘要

被引文献

相似文献

大型数据集对于深度学习是必要的,因为所使用的算法的性能随着数据集大小的增加而增加。糟糕的数据管理实践和建筑行业的低水平数字化是编制大型数据集的一大障碍;在许多情况下,这可能是非常昂贵的。在其他领域,例如计算机视觉,数据增强技术和合成数据已成功用于解决有限数据集的问题。在这项研究中,不完全,稀疏,深度和变分自动编码器作为数据增强和生成合成数据的方法进行了研究。两个地下和架空输电项目的财务数据集作为案例研究。使用自动编码器增强数据集,并使用深度神经网络回归器预测项目成本。所有增强的数据集都比原始数据集产生了更好的结果。平均而言,自动编码器分别为地下和架空数据集提供了7.2%和11.5%的模型得分改进。MAE和RMSE对于所有的自动编码器都是较低的。地下和高架数据集的平均误差改善分别为22.9%和56.5%。变分自动编码器提供了更强大的结果,更好地代表了两个数据集中的属性之间的非线性相关性。本研究的新奇在于,提出了一种改进现有数据集的方法,从而在其他方法不可行时改进深度学习模型的概括性。此外,这项研究为从业者提供了解决大数据集访问受限的方法,从数据中的非线性相关性中提取见解的可视化方法,以及改善数据隐私并使用类似合成数据共享敏感数据的方法。这项研究的主要贡献是,它提出了一种数据增强技术的变换变量数据。已经开发了许多用于变换不变数据的技术,这些技术有助于提高深度学习模型的性能。这项研究表明,自动编码器是一个很好的选择,数据增强变换变量数据。
Large datasets are necessary for deep learning as the performance of the algorithms used increases as the size of the dataset increases. Poor data management practices and the low level of digitisation of the construction industry represent a big hurdle to compiling big datasets; which in many cases can be prohibitively expensive. In other fields, such as computer vision, data augmentation techniques and synthetic data have been used successfully to address issues with limited datasets. In this study, undercomplete, sparse, deep and variational autoencoders are investigated as methods for data augmentation and generation of synthetic data. Two financial datasets of underground and overhead power transmission projects are used as case studies. The datasets were augmented using the autoencoders, and the project cost was predicted using a deep neural network regressor. All the augmented datasets yielded better results than the original dataset. On average the autoencoders provide a model score improvement of 7.2% and 11.5% for the underground and overhead datasets, respectively. MAE and RMSE are lower for all autoencoders as well. The average error improvement for the underground and overhead datasets is 22.9% and 56.5%, respectively. Variational autoencoders provided more robust results and represented better the non-linear correlations among the attributes in both datasets. The novelty of this study is that presents an approach to improve existing datasets and thus improve the generalisation of deep learning models when other approaches are not feasible. Moreover, this study provides practitioners with methods to address the limited access to big datasets, a visualisation method to extract insights from non-linear correlations in data, and a way to improve data privacy and to enable sharing sensitive data using analogous synthetic data. The main contribution to knowledge of this study is that it presents a data augmentation technique for transformation variant data. Many techniques have been developed for transformation invariant data that contributed to improving the performance of deep learning models. This study showed that autoencoders are a good option for data augmentation for transformation variant data.