Visual voice activity detection with optical flow

Visual voice activity detection with optical flow
复制标题

DOI:
10.1049/iet-ipr.2009.0042
复制
发表时间:
2010-12-01
影响因子:
2.3
通讯作者:
Chambers, J. A.
Chambers, J. A.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Aubrey, A. J.;Hicks, Y. A.;Chambers, J. A.

文献摘要

被引文献

相似文献

目前的语音活动检测方法通常只利用声学信息。因此,由于存在其他声源,如另一个扬声器或非平稳噪声,它们很容易被错误分类。为了解决这个问题,作者提出了一种新的语音活动检测方法,该方法仅使用说话人口腔区域形式的视觉信息。这样的视频信息不受声环境的影响。仿真结果表明,该方法可以获得高百分比的正确沉默检测(CSD)和低百分比的假沉默检测(FSD)。与其他两种视觉语音活动检测器的比较表明,所提出的方法始终更加准确,并且平均产生4%的CSD改进。将该方法应用于先前发表的视听卷积盲源分离算法,以提高说话人的可理解性,从而证实了该方法的有效性。
Current voice activity detection methods generally utilise only acoustic information. Therefore they are susceptible to false classification because of the presence of other acoustic sources such as another speaker or non-stationary noise. To address this issue, the authors propose a new method of voice activity detection using solely visual information in the form of a speaker's mouth region. Such video information is not affected by the acoustic environment. Simulations show that a high percentage correct silence detection (CSD) can be obtained with a low percentage false silence detection (FSD). Comparisons with two other visual voice activity detectors show the proposed method to be consistently more accurate, and on average yields a 4% improvement in CSD. The usefulness of the method is confirmed by applying it to a previously published audio-visual convolutive blind source separation algorithm, to increase the intelligibility of a speaker.