Improved Active Multi-Task Representation Learning via Lasso

Improved Active Multi-Task Representation Learning via Lasso
复制标题

DOI:
10.48550/arxiv.2306.02556
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Yiping Wang;Yifang Chen;Kevin G. Jamieson;S. Du
Yiping Wang;Yifang Chen;Kevin G. Jamieson;S. Du
中科院分区:
其他
文献类型:
--
作者:
Yiping Wang;Yifang Chen;Kevin G. Jamieson;S. Du

文献摘要

被引文献

相似文献

为了充分利用源任务的海量数据,克服目标任务样本的稀缺性,基于多任务预训练的表征学习已成为许多应用中的标准方法。然而,到目前为止,大多数已有的工作都是从纯经验的角度来设计源任务选择策略。最近,CITET{chen2022active}给出了第一个主动多任务表示学习(A-MTRL)算法,该算法能够自适应地对源任务进行采样,并使用L2-正则化的目标源-源相关性参数来证明可以降低总的采样复杂度。但他们的工作在理论上是次优的,在总的源样本复杂性方面,并且在一些需要稀疏训练源任务选择的真实场景中不太实用。在这篇文章中,我们解决了这两个问题。具体地说,我们通过给出基于$u^2的策略的下界,证明了基于L1正则关联的策略的严格优势。当$u^1$未知时,我们提出了一个实用的算法,利用套索程序来估计$\nu^1$。在已知情况下,我们的算法成功地恢复了最优解。除了我们的样本复杂性结果之外,我们还描述了我们的基于$^1$的策略在样本成本敏感的环境中的潜力。最后,我们在真实的计算机视觉数据集上进行了实验,验证了该方法的有效性。
To leverage the copious amount of data from source tasks and overcome the scarcity of the target task samples, representation learning based on multi-task pretraining has become a standard approach in many applications. However, up until now, most existing works design a source task selection strategy from a purely empirical perspective. Recently, \citet{chen2022active} gave the first active multi-task representation learning (A-MTRL) algorithm which adaptively samples from source tasks and can provably reduce the total sample complexity using the L2-regularized-target-source-relevance parameter $\nu^2$. But their work is theoretically suboptimal in terms of total source sample complexity and is less practical in some real-world scenarios where sparse training source task selection is desired. In this paper, we address both issues. Specifically, we show the strict dominance of the L1-regularized-relevance-based ($\nu^1$-based) strategy by giving a lower bound for the $\nu^2$-based strategy. When $\nu^1$ is unknown, we propose a practical algorithm that uses the LASSO program to estimate $\nu^1$. Our algorithm successfully recovers the optimal result in the known case. In addition to our sample complexity results, we also characterize the potential of our $\nu^1$-based strategy in sample-cost-sensitive settings. Finally, we provide experiments on real-world computer vision datasets to illustrate the effectiveness of our proposed method.