Quantifying counts and costs via classification

Quantifying counts and costs via classification
复制标题

DOI:
10.1007/s10618-008-0097-y
复制
发表时间:
2008-10-01
影响因子:
4.8
通讯作者:
Forman, George
Forman, George
中科院分区:
计算机科学3区
文献类型:
--
作者:
Forman, George

文献摘要

被引文献

相似文献

许多业务应用程序跟踪随时间的变化,例如,测量流感事件的每月流行率。在需要分类器来识别相关事件的情况下,不完美的分类准确性可能会导致估计类别流行率的重大偏差。本文定义了机器学习的两个研究挑战。“量化”任务是使用可能具有实质上不同分布的训练集来准确估计测试集中的阳性案例(或类别分布)的数量。“成本量化”变量估计与正类相关联的总成本,其中每个案例都标记有成本属性,例如解决案例的费用。量化与传统的分类研究有着非常不同的实用模式。对于这两种形式的量化,本文描述了各种方法,并使用合适的方法对其进行评估,揭示了当训练数据稀缺时,哪些方法可以提供可靠的估计,测试类分布与训练有很大差异,并且阳性类很少,例如,1%阳性。这些优势可以使量化实用于业务用途,即使在分类准确性很差的情况下。
Many business applications track changes over time, for example, measuring the monthly prevalence of influenza incidents. In situations where a classifier is needed to identify the relevant incidents, imperfect classification accuracy can cause substantial bias in estimating class prevalence. The paper defines two research challenges for machine learning. The 'quantification' task is to accurately estimate the number of positive cases (or class distribution) in a test set, using a training set that may have a substantially different distribution. The 'cost quantification' variant estimates the total cost associated with the positive class, where each case is tagged with a cost attribute, such as the expense to resolve the case. Quantification has a very different utility model from traditional classification research. For both forms of quantification, the paper describes a variety of methods and evaluates them with a suitable methodology, revealing which methods give reliable estimates when training data is scarce, the testing class distribution differs widely from training, and the positive class is rare, e.g., 1% positives. These strengths can make quantification practical for business use, even where classification accuracy is poor.