Techniques for dealing with incomplete data: a tutorial and survey

Techniques for dealing with incomplete data: a tutorial and survey
复制标题

DOI:
10.1007/s10044-014-0411-9
复制
发表时间:
2015-02-01
影响因子:
3.9
通讯作者:
Trentin, Edmondo
Trentin, Edmondo
中科院分区:
计算机科学4区
文献类型:
--
作者:
Aste, Marco;Boninsegna, Massimo;Trentin, Edmondo

文献摘要

被引文献

相似文献

模式识别或机器学习算法的实际应用通常会出现数据部分丢失、被噪声破坏或不完整的情况。尽管如此,在过去十年中,机器学习社区的发展主要集中在学习机器的数学分析上,这使得从业者很难对这个问题的主要方法进行概述。奇怪的是,其结果是,即使是植根于统计的既定方法似乎也早已被遗忘。尽管关于这个主题的相关文献如此广泛,以至于现在不可能详尽无遗地报道,但本文的第一个目标是为读者提供一个重要的调查,或者完全合理的技术,用于处理模式识别,机器学习和不完整数据的密度估计。其次,本文旨在为有兴趣的从业者提供一个可行的教程工具,允许独立的,逐步理解几种方法。我们努力将不同的技术分类如下:(1)启发式方法;(2)统计方法;(3)面向连接主义的技术;(4)其他方法(动力系统,对抗性特征删除等)。
Real-world applications of pattern recognition, or machine learning algorithms, often present situations where the data are partly missing, corrupted by noise, or otherwise incomplete. In spite of that, developments in the machine learning community in the last decade have mostly focused on mathematical analysis of learning machines, making it difficult for practitioners to recollect an overview of major approaches to this issue. Paradoxically, as a consequence, even established methodologies rooted in statistics appear to have long been forgotten. Although the relevant literature on the topic is so wide that no exhaustive coverage is nowadays possible, the first goal of this paper is to provide the reader with a nonetheless significant survey of major, or utterly sound, techniques for dealing with the tasks of pattern recognition, machine learning, and density estimation from incomplete data. Secondly, the paper aims at representing a viable tutorial tool for the interested practitioner, by allowing for self-contained, step-by-step understanding of several approaches. An effort is made to categorize the different techniques as follows: (1) heuristic methods; (2) statistical approaches; (3) connectionist-oriented techniques; (4) other approaches (dynamical systems, adversarial deletion of features, etc.).