Classifying with confidence from incomplete information

Classifying with confidence from incomplete information
复制标题

DOI:
10.5555/2567709.2627671
复制
发表时间:
2013
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Nathan Parrish;H. Anderson;M. Gupta;Dun-Yu Hsiao
Nathan Parrish;H. Anderson;M. Gupta;Dun-Yu Hsiao
中科院分区:
其他
文献类型:
--
作者:
Nathan Parrish;H. Anderson;M. Gupta;Dun-Yu Hsiao

文献摘要

被引文献

相似文献

本文研究了在给定不完全信息的情况下对测试样本进行分类的问题。当测试样本的数据随着时间的推移而收集时,或者当计算分类特征必须产生成本时,这个问题自然会出现。例如,在分布式传感器网络中,只有一小部分传感器可能在某个时间报告了测量结果,并且需要额外的时间、功率和带宽来收集完整的数据以进行分类。一个实际的目标是,一旦有足够的数据可以做出正确的决定,就分配一个类别标签。我们正式通过可靠性的概念,这一目标的概率,标签分配给不完整的数据将是相同的标签分配给完整的数据,我们提出了一种方法来分类不完整的数据,只有当一些可靠性阈值得到满足。我们的方法将完整数据建模为随机变量,其分布取决于当前的不完整数据和(完整)训练数据。该方法不同于标准的插补策略,因为我们的重点是确定分类决策的可靠性,而不仅仅是类标签。我们表明,该方法提供了有用的可靠性估计的一组实验的时间序列数据集,其目标是尽可能早地分类的时间序列,同时仍然保证可靠性阈值得到满足的插补类标签的正确性。
We consider the problem of classifying a test sample given incomplete information. This problem arises naturally when data about a test sample is collected over time, or when costs must be incurred to compute the classification features. For example, in a distributed sensor network only a fraction of the sensors may have reported measurements at a certain time, and additional time, power, and bandwidth is needed to collect the complete data to classify. A practical goal is to assign a class label as soon as enough data is available to make a good decision. We formalize this goal through the notion of reliability--the probability that a label assigned given incomplete data would be the same as the label assigned given the complete data, and we propose a method to classify incomplete data only if some reliability threshold is met. Our approach models the complete data as a random variable whose distribution is dependent on the current incomplete data and the (complete) training data. The method differs from standard imputation strategies in that our focus is on determining the reliability of the classification decision, rather than just the class label. We show that the method provides useful reliability estimates of the correctness of the imputed class labels on a set of experiments on time-series data sets, where the goal is to classify the time-series as early as possible while still guaranteeing that the reliability threshold is met.