Mixture-of-Linear-Experts for Long-term Time Series Forecasting

Mixture-of-Linear-Experts for Long-term Time Series Forecasting
复制标题

DOI:
10.48550/arxiv.2312.06786
复制
发表时间:
2023-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Ronghao Ni;Zinan Lin;Shuaiqi Wang;Giulia Fanti
Ronghao Ni;Zinan Lin;Shuaiqi Wang;Giulia Fanti
中科院分区:
其他
文献类型:
--
作者:
Ronghao Ni;Zinan Lin;Shuaiqi Wang;Giulia Fanti

文献摘要

相似文献

长期时间序列预测(LTSF)的目的是在给定过去值的情况下预测时间序列的未来值。在某些情况下,这个问题的当前最先进技术(SOTA)是通过线性中心模型实现的,其主要特征是线性映射层。然而,由于它们固有的简单性,它们不能使其预测规则适应时间序列模式的周期性变化。为了应对这一挑战,我们提出了一个混合专家风格的线性中心模型的增强,并提出了混合线性专家(MoLE)。MoLE不是训练单个模型,而是训练多个以线性为中心的模型(即,专家)和一个路由器模型,加权和混合他们的输出。虽然整个框架是端到端训练的,但每个专家都学习专门研究特定的时间模式,路由器模型学习自适应地组成专家。实验表明,MoLE在我们评估的78%以上的数据集和设置中减少了以线性为中心的模型(包括DLinear,RLinear和RMLP)的预测误差。通过使用MoLE,现有的线性中心模型可以在PatchTST报告的68%的实验中实现SOTA LTSF结果,而现有的单头线性中心模型仅在25%的情况下实现SOTA结果。
Long-term time series forecasting (LTSF) aims to predict future values of a time series given the past values. The current state-of-the-art (SOTA) on this problem is attained in some cases by linear-centric models, which primarily feature a linear mapping layer. However, due to their inherent simplicity, they are not able to adapt their prediction rules to periodic changes in time series patterns. To address this challenge, we propose a Mixture-of-Experts-style augmentation for linear-centric models and propose Mixture-of-Linear-Experts (MoLE). Instead of training a single model, MoLE trains multiple linear-centric models (i.e., experts) and a router model that weighs and mixes their outputs. While the entire framework is trained end-to-end, each expert learns to specialize in a specific temporal pattern, and the router model learns to compose the experts adaptively. Experiments show that MoLE reduces forecasting error of linear-centric models, including DLinear, RLinear, and RMLP, in over 78% of the datasets and settings we evaluated. By using MoLE existing linear-centric models can achieve SOTA LTSF results in 68% of the experiments that PatchTST reports and we compare to, whereas existing single-head linear-centric models achieve SOTA results in only 25% of cases.