Missing Data Imputation Through the Use of the Random Forest Algorithm

Missing Data Imputation Through the Use of the Random Forest Algorithm
复制标题

通过使用随机森林算法进行缺失数据插补

DOI:
10.1007/978-3-642-03156-4_6
复制
发表时间:
2009
影响因子:
--
通讯作者:
T. Marwala
T. Marwala
中科院分区:
--
文献类型:
--
作者:
Adam Pantanowitz;T. Marwala

文献摘要

被引文献

相似文献

本文对缺失数据填补的不同范式进行了比较。使用的数据集是2001年进行的产前诊所研究调查的艾滋病毒血清阳性率数据。数据插补是通过五种方法:随机森林;自联想神经网络与遗传算法;自联想神经模糊配置;和两个随机森林和神经网络的混合动力车。结果表明,随机森林是上级在填补缺失数据的给定数据集的准确性和计算时间方面,与自联想网络相比,某些变量的准确性平均提高高达32%。虽然混合动力系统的概念有希望,所提出的系统似乎是由他们的自联想神经网络组件的阻碍。
This paper presents a comparison of different paradigms used for missing data imputation. The data set used is HIV seroprevalence data from an antenatal clinic study survey performed in 2001. Data imputation is performed through five methods: Random Forests; auto-associative neural networks with genetic algorithms; auto-associative neuro-fuzzy configurations; and two random forest and neural network based hybrids. Results indicate that Random Forests are superior in imputing missing data for the given data set in terms of accuracy and in terms of computation time, with accuracy increases of up to 32 % on average for certain variables when compared with auto-associative networks. While the concept of hybrid systems has promise, the presented systems appear to be hindered by their auto-associative neural network components.