Assisted Learning for Organizations with Limited Imbalanced Data

Assisted Learning for Organizations with Limited Imbalanced Data
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
Trans. Mach. Learn. Res.
影响因子:
--
通讯作者:
Cheng Chen;Jiaying Zhou;Jie Ding;Yi Zhou
Cheng Chen;Jiaying Zhou;Jie Ding;Yi Zhou
中科院分区:
其他
文献类型:
--
作者:
Cheng Chen;Jiaying Zhou;Jie Ding;Yi Zhou

文献摘要

相似文献

在大数据时代,许多大型组织正在将机器学习集成到他们的工作管道中,以促进数据分析。然而,他们训练的模型的性能往往受到有限和不平衡的数据的限制。在这项工作中,我们开发了一个辅助学习框架,以帮助组织提高他们的学习绩效。这些组织有足够的计算资源,但受到严格的数据共享和协作政策的约束。他们有限的不平衡数据往往会导致有偏见的推理和次优决策。在辅助学习中,组织学习者从外部服务提供者购买辅助服务,并旨在仅在几轮辅助中提高其模型性能。我们为辅助深度学习和辅助强化学习开发了有效的随机训练算法。与现有的需要频繁传输梯度或模型的分布式算法不同,我们的框架允许学习者只偶尔与服务提供者共享信息,但仍然可以获得一个模型,该模型可以实现接近Oracle的性能,就好像所有数据都是集中的一样。
In the era of big data, many big organizations are integrating machine learning into their work pipelines to facilitate data analysis. However, the performance of their trained models is often restricted by limited and imbalanced data available to them. In this work, we develop an assisted learning framework for assisting organizations to improve their learning performance. The organizations have sufficient computation resources but are subject to stringent data-sharing and collaboration policies. Their limited imbalanced data often cause biased inference and sub-optimal decision-making. In assisted learning, an organizational learner purchases assistance service from an external service provider and aims to enhance its model performance within only a few assistance rounds. We develop effective stochastic training algorithms for both assisted deep learning and assisted reinforcement learning. Different from existing distributed algorithms that need to frequently transmit gradients or models, our framework allows the learner to only occasionally share information with the service provider, but still obtain a model that achieves near-oracle performance as if all the data were centralized.