Cost-Sensitive Semi-Supervised Support Vector Machine

Cost-Sensitive Semi-Supervised Support Vector Machine
复制标题

DOI:
10.1609/aaai.v24i1.7661
复制
发表时间:
2010-07
期刊:
--
影响因子:
--
通讯作者:
Yu-Feng Li;James T. Kwok;Zhi-Hua Zhou
Yu-Feng Li;James T. Kwok;Zhi-Hua Zhou
中科院分区:
其他
文献类型:
--
作者:
Yu-Feng Li;James T. Kwok;Zhi-Hua Zhou

文献摘要

被引文献

相似文献

在本文中,我们研究了代价敏感的半监督学习,其中许多训练样本是未标记的,并且不同的错误分类错误与不同的代价相关。这种情况在许多实际应用程序中都会发生。例如,在一些疾病诊断中,错误地将患者诊断为健康的成本远远高于将健康的人诊断为患者的成本。此外,标签数据的获取需要昂贵的医疗诊断,而收集基本健康信息等非标签数据要便宜得多。我们提出了代价敏感的半监督支持向量机(CS4VM)来解决这个问题。我们表明,当给定未标记数据的标签均值时,CS4VM与可访问所有未标记数据的地面真实标签的有监督代价敏感支持向量机非常接近。这一观察结果导致了一种高效的算法,该算法首先估计标签均值,然后通过高效的支持向量机求解器用插件标签均值训练CS4VM。在广泛的数据集上的实验表明,该方法能够降低总代价,并且计算效率高。
In this paper, we study cost-sensitive semi-supervised learning where many of the training examples are unlabeled and different misclassification errors are associated with unequal costs. This scenario occurs in many real-world applications. For example, in some disease diagnosis, the cost of erroneously diagnosing a patient as healthy is much higher than that of diagnosing a healthy person as a patient. Also, the acquisition of labeled data requires medical diagnosis which is expensive, while the collection of unlabeled data such as basic health information is much cheaper. We propose the CS4VM (Cost-Sensitive Semi-Supervised Support Vector Machine) to address this problem. We show that the CS4VM, when given the label means of the unlabeled data, closely approximates the supervised cost-sensitive SVM that has access to the ground-truth labels of all the unlabeled data. This observation leads to an efficient algorithm which first estimates the label means and then trains the CS4VM with the plug-in label means by an efficient SVM solver. Experiments on a broad range of data sets show that the proposed method is capable of reducing the total cost and is computationally efficient.