Photograph aesthetical evaluation and classification with deep convolutional neural networks

Photograph aesthetical evaluation and classification with deep convolutional neural networks
复制标题

DOI:
10.1016/j.neucom.2016.08.098
复制
发表时间:
2017-03
期刊:
影响因子:
6
通讯作者:
Yunlan Tan;Pengjie Tang;Yimin Zhou;Wenlang Luo;Yongping Kang;Guangyao Li
Yunlan Tan;Pengjie Tang;Yimin Zhou;Wenlang Luo;Yongping Kang;Guangyao Li
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yunlan Tan;Pengjie Tang;Yimin Zhou;Wenlang Luo;Yongping Kang;Guangyao Li

文献摘要

被引文献

相似文献

为了应对数码摄影及其众多相关应用的发展,研究人员一直在积极研究提供照片的自动化美学评价和分类的方法。为了使计算网络能够识别美学特征,必须使用具有已知美学价值的特征样本集来训练学习算法。开发这种训练的传统方法需要手动提取美学特征,以便在实践数据集中使用。随着卷积神经网络(CNN)的大量出现,网络实现了特征的自动学习,成为评价和分类的重要工具。在我们研究的时候,现有的几种用于照片审美分类的卷积神经网络都是使用浅层网络,这限制了性能的提高。此外,大多数方法只提取一个块作为训练样本,例如从每幅图像中提取一个缩小尺寸的裁剪。然而,单个补丁可能不能准确地表示整个图像,这可能会导致训练过程中的模糊性。更重要的是,对于现有的数据集,每个类别的高质量图像的数量大多太少,无法训练深入的CNN网络。为了解决这些问题,我们引入了一种新颖的具有深度和广度CNN的照片美学分类器来进行细粒度的美学质量预测。首先,我们从DPChallenge.com(一家知名的在线摄影门户网站)下载了大量的消费者摄影图像,以构建适合审美质量评估的数据集。然后,通过双线性插值法将图像放大到256×256,并裁剪10个面片(中心+四角+翻转)。一旦我们将集合与图像的训练标签相关联,我们就将带有补丁的图像馈送到微调网络中。我们提出的计算方法被配置为将照片划分为高审美价值和低审美价值。指定输出(0,1)的训练模式表示对应的图像属于低美学质量集合。同样地,输出为(1,0)的训练模式表示对应的图像属于高美学质量集合。实验结果表明,该方法的分类正确率大于87.10%,明显优于现有的分类方法。此外,我们的实验表明,我们的结果与人类的视觉感知和审美判断基本一致。
In response to the growth of digital photography and its many related applications, researchers have been actively investigating methods for providing automated aesthetical evaluation and classification of photographs. For computational networks to recognize aesthetic qualities, the learning algorithms must be trained using sample sets of characteristics that have known aesthetic values. Traditional methods for developing this training have required manual extraction of aesthetic features for use in the practice datasets. With abundant appearance of convolutional neural networks (CNN), the networks have learned features automatically and have acted as important tools for evaluation and classification. At the time of our research, several existing convolutional neural networks for photograph aesthetical classification only used shallow depth networks, which limit the improvement of performance. In addition, most methods have extracted only one patch as a training sample, such as a down-sized crop from each image. However, a single patch might not represent the entire image accurately, which could cause ambiguity during training. What's more, for existing datasets, the numbers of high quality images of each category are mostly too small to train deep CNN networks. To solve these problems, we introduce a novel photograph aesthetic classifier with a deep and wide CNN for fine-granularity aesthetical quality prediction. First, we download a large number of consumer photographic images from DPChallenge.com (a well-known online photography portal) to construct a dataset suitable for aesthetic quality assessment. Then, we zoom out the images into 256×256 by bilinear interpolation and crop 10 patches (Center+four Corners+Flipping). Once we have associated the set with the image's training labels, we feed the images with the bag of patches into the fine-tuned networks. Our proposed computational method is configured to classify the photographs into high and low aesthetic values. A training pattern specifying an output of (0, 1) indicates that the corresponding image belongs to the “low aesthetic quality” set. Likewise, a training pattern with an output of (1, 0) indicates that the corresponding image belongs to the “high aesthetic quality” set. Experimental results show that the accuracy of classification provided by our method is greater than 87.10%, which is noticeably better than the state-of-the-art methods. In addition, our experiments show that our results are fundamentally consistent with human visual perception and aesthetic judgments.