SynBench: Task-Agnostic Benchmarking of Pretrained Representations using Synthetic Data

SynBench: Task-Agnostic Benchmarking of Pretrained Representations using Synthetic Data
复制标题

DOI:
10.48550/arxiv.2210.02989
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Ching-Yun Ko;Pin-Yu Chen;Jeet Mohapatra;Payel Das;Lucani E. Daniel
Ching-Yun Ko;Pin-Yu Chen;Jeet Mohapatra;Payel Das;Lucani E. Daniel
中科院分区:
其他
文献类型:
--
作者:
Ching-Yun Ko;Pin-Yu Chen;Jeet Mohapatra;Payel Das;Lucani E. Daniel

文献摘要

相似文献

最近在对下游任务的大型模型进行微调方面取得的成功,使深度学习从以任务为中心的模型设计转向与任务无关的表征学习和具体任务的微调。由于预先训练模型的表示被用作不同下游任务的基础,因此提出了一种新的任务不可知框架-使用合成数据来衡量预先训练的模型表示的质量。我们通过理论推导类条件高斯混合的稳健性和精度之间的折衷来建立一个参考。给定一个预先训练的模型,从高斯混合合成的数据的表示被用来与我们的参考文献进行比较,以推断质量。通过比较原始数据和其表示之间的曲线下面积比率,SynBtch为稳健性-准确性性能基准提供了一个可量化的分数。我们的框架适用于接受连续数据输入的广泛的预训练模型,并且独立于下游任务和数据集。实验结果表明,当对下游任务进行微调时,我们的SynBuchch分数与预训练模型的实际线性探测性能很好地吻合。此外,我们的框架可以用来指导预先训练的表示上的稳健线性探测的设计,以缓解下游任务的稳健性和精确度之间的权衡。
Recent success in fine-tuning large models, that are pretrained on broad data at scale, on downstream tasks has led to a significant paradigm shift in deep learning, from task-centric model design to task-agnostic representation learning and task-specific fine-tuning. As the representations of pretrained models are used as a foundation for different downstream tasks, this paper proposes a new task-agnostic framework, \textit{SynBench}, to measure the quality of pretrained representations using synthetic data. We set up a reference by a theoretically-derived robustness-accuracy tradeoff of the class conditional Gaussian mixture. Given a pretrained model, the representations of data synthesized from the Gaussian mixture are used to compare with our reference to infer the quality. By comparing the ratio of area-under-curve between the raw data and their representations, SynBench offers a quantifiable score for robustness-accuracy performance benchmarking. Our framework applies to a wide range of pretrained models taking continuous data inputs and is independent of the downstream tasks and datasets. Evaluated with several pretrained vision transformer models, the experimental results show that our SynBench score well matches the actual linear probing performance of the pre-trained model when fine-tuned on downstream tasks. Moreover, our framework can be used to inform the design of robust linear probing on pretrained representations to mitigate the robustness-accuracy tradeoff in downstream tasks.