Video analysis for the detection of animals using convolutional neural networks and consumer-grade drones

Video analysis for the detection of animals using convolutional neural networks and consumer-grade drones
复制标题

DOI:
10.1139/juvs-2020-0018
复制
发表时间:
2021-06-01
影响因子:
2.3
通讯作者:
Wich, Serge A.
Wich, Serge A.
中科院分区:
其他
文献类型:
--
作者:
Chalmers, C.;Fergus, P.;Wich, Serge A.

文献摘要

被引文献

相似文献

确定动物的分布和密度在保护中很重要。这一过程既费时又费力。无人机已被用于帮助减轻人力密集型任务,在更短的时间内覆盖大的地理区域。在本文中,我们进一步研究了这个想法,使用概念验证来从无人机镜头中检测犀牛和汽车。概念验证利用了现成的技术和消费级无人机硬件。该研究证明了使用机器学习(ML)自动执行常规保护任务的可行性,例如动物检测和跟踪。该原型是使用DJI Mavic Pro 2开发的,并在全球移动的通信系统(GSM)网络上进行了测试。Faster-RCNN Resnet 101架构用于迁移学习。使用帧采样技术执行推断,以解决精度、处理速度和实时视频馈送同步之间所需的权衡。推理模型托管在Web平台上,来自无人机的视频流(使用OcuSync)传输到实时消息传递协议(RTMP)服务器,用于后续分类。在训练过程中,最佳模型的平均精度(mAP)分别为0.83交集大于并集(@IOU)0.50和0.69@IOU 0.75。在诺斯利Safari中测试系统时,我们的原型能够实现以下目标:灵敏度(Sen),0.91(0.869,0.94);特异性(Spec),0.78(0.74,0.82);和准确度(ACC),0.84(0.81,0.87),当检测犀牛时,Sen,1.00(1.00,1.00); Spec,1.00(1.00,1.00);当检测汽车时,ACC,1.00(1.00,1.00)。
Determining animal distribution and density is important in conservation. The process is both time-consuming and labour-intensive. Drones have been used to help mitigate human-intensive tasks by covering large geographical areas over a much shorter timescale. In this paper we investigate this idea further using a proof of concept to detect rhinos and cars from drone footage. The proof of concept utilises off-the-shelf technology and consumer-grade drone hardware. The study demonstrates the feasibility of using machine learning (ML) to automate routine conservation tasks, such as animal detection and tracking. The prototype has been developed using a DJI Mavic Pro 2 and tested over a global system for mobile communications (GSM) network. The Faster-RCNN Resnet 101 architecture is used for transfer learning. Inference is performed with a frame sampling technique to address the required trade-off between precision, processing speed, and live video feed synchronisation. Inference models are hosted on a web platform and video streams from the drone (using OcuSync) are transmitted to a real-time messaging protocol (RTMP) server for subsequent classification. During training, the best model achieves a mean average precision (mAP) of 0.83 intersection over union (@IOU) 0.50 and 0.69 @IOU 0.75, respectively. On testing the system in Knowsley Safari our prototype was able to achieve the following: sensitivity (Sen), 0.91 (0.869, 0.94); specificity (Spec), 0.78 (0.74, 0.82); and an accuracy (ACC), 0.84 (0.81, 0.87) when detecting rhinos, and Sen, 1.00 (1.00, 1.00); Spec, 1.00 (1.00, 1.00); and an ACC, 1.00 (1.00, 1.00) when detecting cars.