How Important is the Train-Validation Split in Meta-Learning?

How Important is the Train-Validation Split in Meta-Learning?
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yu Bai;Minshuo Chen;Pan Zhou;T. Zhao;J. Lee;S. Kakade;Haiquan Wang;Caiming Xiong
Yu Bai;Minshuo Chen;Pan Zhou;T. Zhao;J. Lee;S. Kakade;Haiquan Wang;Caiming Xiong
中科院分区:
其他
文献类型:
--
作者:
Yu Bai;Minshuo Chen;Pan Zhou;T. Zhao;J. Lee;S. Kakade;Haiquan Wang;Caiming Xiong

文献摘要

被引文献

相似文献

元学习的目标是通过从多个现有任务中学习一个“先验”来对新任务进行快速适应。元学习中的一个常见做法是执行训练验证分割,其中先验在数据的一个分割上适应任务,并且在另一个分割上评估结果预测器。尽管它很流行,但无论是在理论上还是在实践中,训练-验证分割的重要性都没有得到很好的理解,特别是与更直接的非分割方法相比,该方法使用所有每个任务的数据进行训练和评估。我们提供了一个详细的理论研究,在任务数量趋于无穷大的渐近设置中,训练验证分裂是否以及何时有助于线性质心元学习问题。我们表明,分裂方法如预期地收敛到最佳先验,而非分裂方法在没有对数据进行结构性假设的情况下通常不会收敛到最佳先验。相比之下,如果数据是从线性模型(可实现的制度),我们表明,分裂和非分裂的方法收敛到最佳的先验。此外,也许令人惊讶的是,我们的主要结果表明,非分裂方法在这种数据分布下实现了严格更好的渐近超额风险,即使在正则化参数和分裂比为两种方法进行了优化调整。我们的研究结果强调,数据分割可能并不总是可取的,特别是当数据是可实现的模型。我们通过实验验证了我们的理论,表明非分裂方法确实可以在模拟和真实的元学习任务上优于分裂方法。
Meta-learning aims to perform fast adaptation on a new task through learning a "prior" from multiple existing tasks. A common practice in meta-learning is to perform a train-validation split where the prior adapts to the task on one split of the data, and the resulting predictor is evaluated on another split. Despite its prevalence, the importance of the train-validation split is not well understood either in theory or in practice, particularly in comparison to the more direct non-splitting method, which uses all the per-task data for both training and evaluation. We provide a detailed theoretical study on whether and when the train-validation split is helpful on the linear centroid meta-learning problem, in the asymptotic setting where the number of tasks goes to infinity. We show that the splitting method converges to the optimal prior as expected, whereas the non-splitting method does not in general without structural assumptions on the data. In contrast, if the data are generated from linear models (the realizable regime), we show that both the splitting and non-splitting methods converge to the optimal prior. Further, perhaps surprisingly, our main result shows that the non-splitting method achieves a strictly better asymptotic excess risk under this data distribution, even when the regularization parameter and split ratio are optimally tuned for both methods. Our results highlight that data splitting may not always be preferable, especially when the data is realizable by the model. We validate our theories by experimentally showing that the non-splitting method can indeed outperform the splitting method, on both simulations and real meta-learning tasks.