Missing Data Imputation Through the Use of the Random Forest Algorithm
Missing Data Imputation Through the Use of the Random Forest Algorithm
复制标题
通过使用随机森林算法进行缺失数据插补
DOI:
10.1007/978-3-642-03156-4_6
复制
发表时间:
2009
影响因子:
--
通讯作者:
T. Marwala
中科院分区:
文献类型:
--
作者:
Adam Pantanowitz;T. Marwala
This paper presents a comparison of different paradigms used for missing data imputation. The data set used is HIV seroprevalence data from an antenatal clinic study survey performed in 2001. Data imputation is performed through five methods: Random Forests; auto-associative neural networks with genetic algorithms; auto-associative neuro-fuzzy configurations; and two random forest and neural network based hybrids. Results indicate that Random Forests are superior in imputing missing data for the given data set in terms of accuracy and in terms of computation time, with accuracy increases of up to 32 % on average for certain variables when compared with auto-associative networks. While the concept of hybrid systems has promise, the presented systems appear to be hindered by their auto-associative neural network components.