Learning from positive and unlabeled examples with different data distributions

Learning from positive and unlabeled examples with different data distributions
复制标题

DOI:
10.1007/11564096_24
复制
发表时间:
2005-01-01
期刊:
MACHINE LEARNING: ECML 2005, PROCEEDINGS
影响因子:
--
通讯作者:
Liu, B
Liu, B
中科院分区:
其他
文献类型:
--
作者:
Li, XL;Liu, B

文献摘要

被引文献

相似文献

我们研究的问题,学习从积极的和未标记的例子。虽然有几种技术可以处理这个问题,但它们都假设正集P中的正例和未标记集U中的正例是从相同的分布生成的。这一假设在实践中可能会被违反。例如,想要从Web收集所有打印机页面。可以使用来自一个站点的打印机页面作为正页面的集合P,并且使用来自另一站点的产品页面作为U。我们希望将U中的页面分为打印机页面和非打印机页面。虽然这两个网站的打印机页面有很多相似之处,但它们也可能有很大的不同,因为不同的网站通常以不同的风格呈现类似的产品,并有不同的重点。在这种情况下,现有的方法表现不佳。本文提出了一种新的技术A-EM来处理这个问题。产品页面分类的实验结果证明了该方法的有效性。
We study the problem of learning from positive and unlabeled examples. Although several techniques exist for dealing with this problem, they all assume that positive examples in the positive set P and the positive examples in the unlabeled set U are generated from the same distribution. This assumption may be violated in practice. For example, one wants to collect all printer pages from the Web. One can use the printer pages from one site as the set P of positive pages and use product pages from another site as U. One wants to classify the pages in U into printer pages and non-printer pages. Although printer pages from the two sites have many similarities, they can also be quite different because different sites often present similar products in different styles and have different focuses. In such cases, existing methods perform poorly. This paper proposes a novel technique A-EM to deal with the problem. Experiment results with product page classification demonstrate the effectiveness of the proposed technique.