GANDaLF: GAN for Data-Limited Fingerprinting

GANDaLF: GAN for Data-Limited Fingerprinting
复制标题

DOI:
10.2478/popets-2021-0029
复制
发表时间:
2021-01
影响因子:
--
通讯作者:
Se Eun Oh;Nate Mathews;Mohammad Saidur Rahman;M. Wright;Nicholas Hopper
Se Eun Oh;Nate Mathews;Mohammad Saidur Rahman;M. Wright;Nicholas Hopper
中科院分区:
--
文献类型:
--
作者:
Se Eun Oh;Nate Mathews;Mohammad Saidur Rahman;M. Wright;Nicholas Hopper

文献摘要

被引文献

相似文献

摘要:本文介绍了一种新的基于深度学习的基于生成对抗网络的数据有限指纹识别技术(GANDaLF),用于对Tor流量进行网站指纹识别(WF)。与大多数早期关于WF深度学习的工作相反,GANDaLF旨在使用很少的训练样本,并通过使用生成式对抗网络生成大量“假”数据来实现这一目标,这些数据有助于训练深度神经网络来区分实际训练数据的类别。我们在低数据场景下评估GANDaLF,包括每个站点少至10个训练实例,以及多种设置,包括网站索引页面的指纹识别和网站内非索引页面的指纹识别。在标准WF设置下,每个站点仅20个实例(100个站点),GANDaLF实现了87%的封闭世界准确率。特别是,在子页面指纹识别的所有设置中,GANDaLF都可以优于Var-CNN和三重指纹(TF)。例如,在每个站点使用20个实例的训练集上,GANDaLF比TF高出29%,Var-CNN高出38%。
Abstract We introduce Generative Adversarial Networks for Data-Limited Fingerprinting (GANDaLF), a new deep-learning-based technique to perform Website Fingerprinting (WF) on Tor traffic. In contrast to most earlier work on deep-learning for WF, GANDaLF is intended to work with few training samples, and achieves this goal through the use of a Generative Adversarial Network to generate a large set of “fake” data that helps to train a deep neural network in distinguishing between classes of actual training data. We evaluate GANDaLF in low-data scenarios including as few as 10 training instances per site, and in multiple settings, including fingerprinting of website index pages and fingerprinting of non-index pages within a site. GANDaLF achieves closed-world accuracy of 87% with just 20 instances per site (and 100 sites) in standard WF settings. In particular, GANDaLF can outperform Var-CNN and Triplet Fingerprinting (TF) across all settings in subpage fingerprinting. For example, GANDaLF outperforms TF by a 29% margin and Var-CNN by 38% for training sets using 20 instances per site.