Mix and Match: An Optimistic Tree-Search Approach for Learning Models from Mixture Distributions

Mix and Match: An Optimistic Tree-Search Approach for Learning Models from Mixture Distributions
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Matthew Faw;Rajat Sen;Karthikeyan Shanmugam;C. Caramanis;S. Shakkottai
Matthew Faw;Rajat Sen;Karthikeyan Shanmugam;C. Caramanis;S. Shakkottai
中科院分区:
其他
文献类型:
--
作者:
Matthew Faw;Rajat Sen;Karthikeyan Shanmugam;C. Caramanis;S. Shakkottai

文献摘要

相似文献

我们考虑一个协变量移位问题,其中对于同一个学习问题,可以访问几个不同的训练数据集,并且可以访问一个小的验证集,该验证集可能与所有单独的训练分布不同。这种协变量的变化部分是由于数据集中未观察到的特征造成的。然后,目标是在训练数据集上找到最佳混合分布(仅具有观察到的特征),以便使用此混合来训练学习算法具有最佳验证性能。我们提出的算法,${\sf Mix\&Match}$,结合了随机梯度下降(SGD)与乐观树搜索和模型重用(从不同的混合物分布的样本进化部分训练模型)在混合物的空间,为这项任务。我们证明简单的遗憾保证我们的算法恢复最佳的混合物,鉴于SGD评估的总预算。最后,我们在两个真实世界的数据集上验证了我们的算法。
We consider a covariate shift problem where one has access to several different training datasets for the same learning problem and a small validation set which possibly differs from all the individual training distributions. This covariate shift is caused, in part, due to unobserved features in the datasets. The objective, then, is to find the best mixture distribution over the training datasets (with only observed features) such that training a learning algorithm using this mixture has the best validation performance. Our proposed algorithm, ${\sf Mix\&Match}$, combines stochastic gradient descent (SGD) with optimistic tree search and model re-use (evolving partially trained models with samples from different mixture distributions) over the space of mixtures, for this task. We prove simple regret guarantees for our algorithm with respect to recovering the optimal mixture, given a total budget of SGD evaluations. Finally, we validate our algorithm on two real-world datasets.