Leveraging Visual Supervision for Array-Based Active Speaker Detection and Localization

Leveraging Visual Supervision for Array-Based Active Speaker Detection and Localization
复制标题

DOI:
10.1109/taslp.2023.3346643
复制
发表时间:
2023-12
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Davide Berghi;Philip J. B. Jackson
Davide Berghi;Philip J. B. Jackson
中科院分区:
其他
文献类型:
--
作者:
Davide Berghi;Philip J. B. Jackson

文献摘要

相似文献

用于主动说话者检测(ASD)的传统视听方法通常依赖于视觉上预提取的面部轨迹和对应的单通道音频来找到视频中的说话者。因此,每当说话者的脸不可见时,它们往往会失败。我们证明了一个简单的音频卷积递归神经网络(CRNN)训练的空间输入特征提取多通道音频可以执行同时水平主动扬声器检测和定位(ASDL),独立于视觉模态。为了解决生成地面真值标签来训练这样一个系统的时间和成本问题,我们提出了一种新的自我监督训练管道,它包含一种“学生-教师”学习方法。一个传统的预先训练的主动说话人检测器被采用作为一个“教师”网络,以提供作为伪标签的说话人的位置。训练多通道音频“学生”网络以生成相同的结果。在推理时,学生网络还可以概括和定位教师网络无法视觉检测到的被遮挡的说话者,从而大大提高了召回率。在TragicTalkers数据集上的实验表明,用所提出的自监督学习方法训练的音频网络可以超过典型的视听方法的性能,并产生与昂贵的传统监督训练相竞争的结果。我们证明,可以实现最小的人工监督时,在学习管道的改进。可以利用更大的训练集和将视觉与多通道音频系统集成来寻求进一步的增益。
Conventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend to fail every time the face of the speaker is not visible. We demonstrate that a simple audio convolutional recurrent neural network (CRNN) trained with spatial input features extracted from multichannel audio can perform simultaneous horizontal active speaker detection and localization (ASDL), independently of the visual modality. To address the time and cost of generating ground truth labels to train such a system, we propose a new self-supervised training pipeline that embraces a “student-teacher” learning approach. A conventional pre-trained active speaker detector is adopted as a “teacher” network to provide the position of the speakers as pseudo-labels. The multichannel audio “student” network is trained to generate the same results. At inference, the student network can generalize and locate also the occluded speakers that the teacher network is not able to detect visually, yielding considerable improvements in recall rate. Experiments on the TragicTalkers dataset show that an audio network trained with the proposed self-supervised learning approach can exceed the performance of the typical audio-visual methods and produce results competitive with the costly conventional supervised training. We demonstrate that improvements can be achieved when minimal manual supervision is introduced in the learning pipeline. Further gains may be sought with larger training sets and integrating vision with the multichannel audio system.