Fast Online Value-Maximizing Prediction Sets with Conformal Cost Control

Fast Online Value-Maximizing Prediction Sets with Conformal Cost Control
复制标题

DOI:
10.48550/arxiv.2302.00839
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhen Lin;Shubhendu Trivedi;Cao Xiao;Jimeng Sung
Zhen Lin;Shubhendu Trivedi;Cao Xiao;Jimeng Sung
中科院分区:
其他
文献类型:
--
作者:
Zhen Lin;Shubhendu Trivedi;Cao Xiao;Jimeng Sung

文献摘要

被引文献

相似文献

许多现实世界的多标签预测问题涉及集值预测,这些集值预测必须满足下游使用所规定的特定要求。我们关注的是一个典型的场景,在这个场景中,分别编码$\textit{value}$和$\textit{cost}$的需求会相互竞争。例如,医院可能希望智能诊断系统能够捕获尽可能多的严重疾病,通常是并发疾病(价值),同时严格控制不正确的预测(成本)。我们提出了一个通用的管道,称为FavMac,以最大限度地提高价值,同时控制在这种情况下的成本。FavMac可以与几乎任何多标签分类器相结合,为成本控制提供无分布的理论保证。此外,与之前的作品不同的是,它可以通过精心设计的在线更新机制来处理现实世界的大规模应用程序,这是独立感兴趣的。我们的方法和理论贡献得到了几个医疗保健任务和合成数据集的实验的支持- FavMac与几个变体和基线相比具有更高的价值,同时保持严格的成本控制。我们的代码可在https://github.com/zlin7/FavMac上获得
Many real-world multi-label prediction problems involve set-valued predictions that must satisfy specific requirements dictated by downstream usage. We focus on a typical scenario where such requirements, separately encoding $\textit{value}$ and $\textit{cost}$, compete with each other. For instance, a hospital might expect a smart diagnosis system to capture as many severe, often co-morbid, diseases as possible (the value), while maintaining strict control over incorrect predictions (the cost). We present a general pipeline, dubbed as FavMac, to maximize the value while controlling the cost in such scenarios. FavMac can be combined with almost any multi-label classifier, affording distribution-free theoretical guarantees on cost control. Moreover, unlike prior works, it can handle real-world large-scale applications via a carefully designed online update mechanism, which is of independent interest. Our methodological and theoretical contributions are supported by experiments on several healthcare tasks and synthetic datasets - FavMac furnishes higher value compared with several variants and baselines while maintaining strict cost control. Our code is available at https://github.com/zlin7/FavMac