Non-intrusive Estimation of Packet Loss Rates in Speech Communication Systems Using Convolutional Neural Networks

Non-intrusive Estimation of Packet Loss Rates in Speech Communication Systems Using Convolutional Neural Networks
复制标题

使用卷积神经网络对语音通信系统中的数据包丢失率进行非侵入式估计

DOI:
10.1109/ism.2018.00026
复制
发表时间:
2018
期刊:
2018 IEEE International Symposium on Multimedia (ISM)
影响因子:
--
通讯作者:
Sebastian Möller
Sebastian Möller
中科院分区:
--
文献类型:
--
作者:
Gabriel Mittag;Sebastian Möller

文献摘要

参考文献

被引文献

相似文献

在本文中,我们分析了深度卷积神经网络是否可以用于检测语音通信系统中的丢失数据包。现代通信网络的语音质量最近已经显著提高,例如通过更高的可用音频带宽。除其他原因外,这是通过使用基于分组的网络而可能的,该网络允许从发送者到接收者终端的全数字传输。然而,这些网络经常遭受由于传输错误而丢失分组所引起的频繁中断。因此,分组丢失率是语音通信服务质量的主要指标之一。尽管如此,网络中丢失了多少数据包的信息并不总是可用的。为了估计丢失数据包的数量,我们计算传输的语音信号的频谱图,并将其用作卷积神经网络的输入。这种方法最近在音乐和语音的检测和识别任务领域得到了普及。由丢失的数据包引起的中断通常可以在降级信号的频谱图中清楚地看到。因此,将频谱图解释为图像并使用图像分类常用的深度学习方法似乎很自然。所提出的模型允许通过简单地使用从接收器侧记录的语音文件来估计通信系统的分组丢失率,而不需要最初通过信道发送的参考语音信号。我们的结果表明,与基于MFCC特征的模型相比,该模型将预测误差降低了75%以上。
In this paper, we analyze whether deep convolutional neural networks can be used to detect lost packets in speech communication systems. The speech quality of modern communication networks has significantly improved recently, for example through higher available audio bandwidth. This was, among other reasons, possible through the use of packet-based networks, which allow a fully digital transmission from the sender to the receiver terminal. However, these networks often suffer from frequent interruptions caused by lost packets due to transmission errors. Consequently, the packet loss rate is one of the main indicators for the quality of speech communication services. In spite of that, the information of how many packets are lost in a network is not always available. To estimate the amount of lost packets, we calculate spectrograms of the transmitted speech signals and use them as input of a convolutional neural network. This approach has recently gained popularity in the field of detection and recognition tasks for music and speech. The interruptions caused by lost packets can often clearly be seen in the spectrogram of the degraded signal. Therefore, it seems natural to interpret the spectrograms as images and use deep learning methods that are common for image classification. The proposed model allows for estimating the packet loss rate of a communication system by simply using the recorded speech file from the receiver side, without the need of the reference speech signal that was originally sent through the channel. Our results show that the model reduces the prediction error by more than 75% when compared to a model that is based on MFCC features.
AMR-WB 编解码器的比特率和串联检测及其在网络测试中的应用
DOI: 10.23919/eusipco.2018.8553360
发表时间: 2018
期刊: 2018 26th European Signal Processing Conference (EUSIPCO)
影响因子: --
作者:
T. Hübschen;G. Schmidt
通讯作者: G. Schmidt
使用共振峰特征和决策树学习检测数据包丢失隐藏
DOI: 10.21437/interspeech.2018-1098
发表时间: 2018
期刊:
影响因子: --
作者:
G. Mittag;S. Möller
通讯作者: S. Möller