A multi-tasking model of speaker-keyword classification for keeping human in the loop of drone-assisted inspection

A multi-tasking model of speaker-keyword classification for keeping human in the loop of drone-assisted inspection
复制标题

说话者关键词分类的多任务模型,使人类能够参与无人机辅助检查的循环

DOI:
10.1016/j.engappai.2022.105597
复制
发表时间:
2023
影响因子:
8
通讯作者:
Qin, Ruwen
Qin, Ruwen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Yu;Parsan, Anisha;Wang, Bill;Dong, Penghao;Yao, Shanshan;Qin, Ruwen

文献摘要

相似文献

音频命令是一种首选的通信媒介,可以让检查人员保持在半自动无人机执行的民用基础设施检查的循环中。为了理解来自一组异构和动态检查员的特定于作业的命令,必须为该组开发一个具有成本效益的模型,并且当该组发生变化时可以轻松地进行调整。本文旨在构建一个具有共享-拆分-协作架构的多任务深度学习模型。这种架构允许两个分类任务共享特征提取器,然后通过特征投影和协作训练来分割在提取的特征中交织的主题特定和关键字特定特征。一组五个授权的主题的基础模型进行训练和测试的检查关键字数据集收集本研究。该模型在对任何授权检查员的关键词进行分类时,平均准确率达到95.3%或更高。该方法在说话人分类中的平均准确率为99.2%。由于模型从池化训练数据中学习到的关键字表示更丰富,因此将基础模型适配到新的检查员只需要来自该检查员的少量训练数据,例如每个关键字五个话语。使用说话人分类分数进行检查员验证,可以实现至少93.9%的成功率在验证授权检查员和76.1%的检测未授权的。此外,该文件证明了该模型的适用性,较大规模的群体在一个公共数据集。本文提供了一种解决方案,以解决人工智能辅助人机交互所面临的挑战,包括工人异质性工人动态和工作异质性。
Audio commands are a preferred communication medium to keep inspectors in the loop of civil infrastructure inspection performed by a semi-autonomous drone. To understand job-specific commands from a group of heterogeneous and dynamic inspectors, a model must be developed cost-effectively for the group and easily adapted when the group changes. This paper is motivated to build a multi-tasking deep learning model that possesses a Share–Split–Collaborate architecture. This architecture allows the two classification tasks to share the feature extractor and then split subject-specific and keyword-specific features intertwined in the extracted features through feature projection and collaborative training. A base model for a group of five authorized subjects is trained and tested on the inspection keyword dataset collected by this study. The model achieved a 95.3% or higher mean accuracy in classifying the keywords of any authorized inspectors. Its mean accuracy in speaker classification is 99.2%. Due to the richer keyword representations that the model learns from the pooled training data Adapting the base model to a new inspector requires only a little training data from that inspector Like five utterances per keyword. Using the speaker classification scores for inspector verification can achieve a success rate of at least 93.9% in verifying authorized inspectors and 76.1% in detecting unauthorized ones. Further The paper demonstrates the applicability of the proposed model to larger-size groups on a public dataset. This paper provides a solution to addressing challenges facing AI-assisted human–robot interaction Including worker heterogeneity Worker dynamics And job heterogeneity.