Artificial intelligence technologies, Image and vision computing, Digital signal processing
Artificial intelligence technologies, Image and vision computing, Digital signal processing
批准号:
2124799
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
自动编码器是在机器学习的许多领域中流行的神经网络。它们由两个主要部分组成--编码器和解码器。前者将输入数据转换为向量表示形式,后者从中恢复数据。自动编码器的主要目标是将输入数据和恢复数据之间的差异降至最低。同样重要的是潜在表示,它是编码器输出的矢量。该向量包含重建数据所需的基本低维信息,因此自动编码器通常用于数据压缩。然而,在许多应用中,不需要恢复原始数据,例如某些计算机视觉任务,包括图像分类或检索,其中只需要图像的表示来从图像中进行准确的预测,所提出的研究的重点是在应用场景施加的极端约束下的图像检索,例如在具有噪声的通信信道中,在有限的时间内只能传输少量数据。将研究深度学习(基于深度神经网络)方法来执行这些任务,并且约束以及噪声模型将被纳入系统中。将考虑两个具体的应用场景:人员重新识别和人脸识别闭路电视在需要高效访问大量数据的监控任务中,因此存在内存和噪声限制。人脸识别是一项类似于个人重新识别的任务,但成功识别所需编码的细节类型不同。将测试各种方法来执行这些任务。作为基准,将使用传统的图像压缩方法来满足信道带宽限制,并且将对恢复(解压缩)的图像执行检索任务。随后,压缩算法将被自动编码器取代,在自动编码器中模拟潜在表示的传输,由解码器重建的图像将被送入检索网络。替代方法将包括将表示编码器应用于原始图像并模拟所产生的特征向量的传输。编码方法的有效性将通过识别方法的性能和编码表示在不同量噪声下的紧凑性来评估,本研究的主要目的是研究深度学习自动编码器模型的噪声、检索性能和表示压缩之间的关系,并将其与仅基于发送特征向量的替代方法进行比较。在有噪声的通信信道下,编码特定的、面向任务的图像表示的问题尚未被研究。将噪声通信信道模型融入到其设计中的自动编码器是相对较新的,通常使用标准度量来评估,例如MSE和PSNR(脉冲信噪比),而本研究将考虑面向任务从而更可靠的度量,即检索性能。该研究与许多监控场景相关,其中需要在有限的时间内通过无线通信信道传输大量数据。这类应用的一个例子是使用配备摄像头的无人机在一大群人中搜索嫌疑人。这一应用带来的限制,例如使用低功率发射机使传输的信号容易受到噪声的影响,以及传输的信息的高度压缩水平,对通常为身份识别任务提出的方法提出了新的挑战。它们需要新的编码方法来编码图像表示,该方法对噪声具有健壮性,并且面向任务,以仅保留与任务相关的信息。
英文摘要
Autoencoders are neural networks popular in many areas of Machine Learning. They are built of two main parts - encoder and decoder. The former transforms input data into a vector representation and the latter recovers the data from it. The main objective of the autoencoders is to minimize difference between the input and the recovered data. Equally important is the latent representation, which is a vector output by the encoder. This vector contains essential, low-dimensional information needed to reconstruct the data, therefore autoencoders are often used for data compression. However, there are many applications, where recovering the original data is not necessary such as certain Computer Vision tasks, including image classification or retrieval, where only the image representation is needed to enable accurate prediction from the image.The focus of the proposed research is on image retrieval under extreme constraints imposed by the application scenario e.g. communication channel with noise where only a small amount of data can be transmitted in a limited time. Deep learning (based on deep neural networks) methods will be investigated to perform these tasks, and the constraints as well as a noise model will be incorporated into the system. Two specific application scenarios will be considered: person re-identification and face recognition CCTV in surveillance tasks where large volumes of data need to be accessed efficiently, hence the memory and noise constraints. Face recognition is a similar task to person re-identification, but the type of details needed to be encoded for successful identification differ.Various approaches will be tested to perform these tasks. As a baseline, conventional image compression methods will be used to meet channel bandwidth limitations and retrieval tasks will be performed on the recovered (decompressed) images. Subsequently, compression algorithms will be replaced by autoencoder, where latent representation's transmission will be simulated, and images reconstructed by the decoder will be fed into the retrieval network. Alternative approach will consist of applying the representation encoders to the original images and simulating transmission of the resulting feature vector. The effectiveness of the encoding approach will be assessed by the performance of the identification methods and the compactness of the encoded representation while subjected to various amounts of noise.The main objective of this research will be to study the relation between noise, retrieval performance, and representation compression for deep learning autoencoder models and compare it to the alternative approach based on sending feature vectors only. The problem of encoding specific and task-oriented image representation under noisy communication channel has not been investigated yet. Autoencoders that incorporate noisy communication channel models into their designs are relatively new and typically evaluated with standard metrics, such as MSE and PSNR (Pulse Signal to Noise Ratio), whereas this research will consider task-oriented thus more reliable metrics, namely retrieval performances.This research is relevant to numerous surveillance scenarios, where it is necessary to transmit large amounts of data via wireless communication channel within a limited time. An example of such application is using a drone equipped with a video camera and searching for a suspect in a large crowd. The constraints introduced by this application such as the use of low power transmitters making the transmitted signal vulnerable to noise, and high compression level of the transmitted information pose new challenges on the methods typically proposed for the human identification tasks. They require new methods of encoding image representations that are robust to noise and task-oriented for retaining only the task relevant information.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
跨文化团队中团队协调机制和团队效能的研究:文化智力的视角
-
批准号:71072055
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2010
-
负责人:唐宁玉
-
依托单位:
基于混沌动力学与复杂网络的群智能优化研究
-
批准号:60673098
-
项目类别:面上项目
-
资助金额:26.0万元
-
批准年份:2006
-
负责人:杨义先
-
依托单位:
智力超常儿童的基因分型的初步研究
-
批准号:30670716
-
项目类别:面上项目
-
资助金额:30.0万元
-
批准年份:2006
-
负责人:施建农
-
依托单位: