Classification without labels: learning from mixed samples in high energy physics

Classification without labels: learning from mixed samples in high energy physics
复制标题

DOI:
10.1007/jhep10(2017)174
复制
发表时间:
2017-08
影响因子:
5.4
通讯作者:
E. Metodiev;B. Nachman;J. Thaler
E. Metodiev;B. Nachman;J. Thaler
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
E. Metodiev;B. Nachman;J. Thaler

文献摘要

被引文献

相似文献

现代机器学习技术可以用来构建强大的模型,解决对撞机物理难题。然而,在许多应用中,由于数据中缺乏真实水平的信息,这些模型是在不完美的模拟上训练的,这可能会导致模拟的模型学习伪像。在本文中,我们介绍了无标签分类(CWoLa)的范例,其中训练分类器以区分类的统计混合物,这在对撞机物理学中很常见。至关重要的是,既不需要单独的标签,也不需要类的比例,但我们证明了CWoLa范式中的最佳分类器也是传统的全监督情况下的最佳分类器,其中所有的标签信息都是可用的。在一个分析玩具例子中展示了这种方法的强大功能之后,我们考虑了对撞机物理学的一个现实基准:使用混合夸克/胶子训练样本区分夸克与胶子引发的喷流。更一般地说,CWoLa可以应用于任何分类问题,其中标签或类别比例是未知的,或者模拟是不可靠的,但是类别的统计混合是可用的。
Modern machine learning techniques can be used to construct powerful models for difficult collider physics problems. In many applications, however, these models are trained on imperfect simulations due to a lack of truth-level information in the data, which risks the model learning artifacts of the simulation. In this paper, we introduce the paradigm of classification without labels (CWoLa) in which a classifier is trained to distinguish statistical mixtures of classes, which are common in collider physics. Crucially, neither individual labels nor class proportions are required, yet we prove that the optimal classifier in the CWoLa paradigm is also the optimal classifier in the traditional fully-supervised case where all label information is available. After demonstrating the power of this method in an analytical toy example, we consider a realistic benchmark for collider physics: distinguishing quark-versus gluon-initiated jets using mixed quark/gluon training samples. More generally, CWoLa can be applied to any classification problem where labels or class proportions are unknown or simulations are unreliable, but statistical mixtures of the classes are available.