Improving Fractal Pre-training

Improving Fractal Pre-training
复制标题

DOI:
10.1109/wacv51458.2022.00247
复制
发表时间:
2021-10
期刊:
2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Connor Anderson;Ryan Farrell
Connor Anderson;Ryan Farrell
中科院分区:
其他
文献类型:
--
作者:
Connor Anderson;Ryan Farrell

文献摘要

被引文献

相似文献

现代计算机视觉系统中使用的深度神经网络需要大量的图像数据集来训练它们。这些精心策划的数据集通常有一百万或更多的图像,跨越一千或更多不同的类别。创建和管理这样一个数据集的过程是一项艰巨的任务,需要付出大量的努力和标签费用,并需要仔细导航标签准确性,版权所有权和内容偏见等技术和社会问题。如果我们有一种方法来利用大型图像数据集的力量,但很少或根本没有目前面临的主要问题和担忧,会怎么样?本文扩展了Kataoka等人最近的工作。[15],提出了一种基于动态生成的分形图像的改进的预训练数据集。大规模图像数据集的存储问题成为分形预训练的优雅之处:零成本的完美标签准确性;无需存储/传输大型图像档案;没有隐私/人口统计偏见/不适当内容的担忧,因为没有人被描绘;图像的无限供应和多样性;图像是免费/开源的。也许令人惊讶的是,避免这些困难只会对性能造成很小的影响。利用新提出的预训练任务-多实例预测-我们的实验表明,对使用分形预训练的网络进行微调,可以达到ImageNet预训练网络的92.7-98.1%的准确率。我们的代码是公开的。1
The deep neural networks used in modern computer vision systems require enormous image datasets to train them. These carefully-curated datasets typically have a million or more images, across a thousand or more distinct categories. The process of creating and curating such a dataset is a monumental undertaking, demanding extensive effort and labelling expense and necessitating careful navigation of technical and social issues such as label accuracy, copyright ownership, and content bias.What if we had a way to harness the power of large image datasets but with few or none of the major issues and concerns currently faced? This paper extends the recent work of Kataoka et al. [15], proposing an improved pre-training dataset based on dynamically-generated fractal images. Challenging issues with large-scale image datasets become points of elegance for fractal pre-training: perfect label accuracy at zero cost; no need to store/transmit large image archives; no privacy/demographic bias/concerns of inappropriate content, as no humans are pictured; limitless supply and diversity of images; and the images are free/open-source. Perhaps surprisingly, avoiding these difficulties imposes only a small penalty in performance. Leveraging a newly-proposed pre-training task—multi-instance prediction—our experiments demonstrate that fine-tuning a network pre-trained using fractals attains 92.7-98.1% of the accuracy of an ImageNet pre-trained network. Our code is publicly available.1