Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition

Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition
复制标题

DOI:
10.1109/icassp49357.2023.10096861
复制
发表时间:
2023-06
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Wei-Cheng Lin;C. Busso
Wei-Cheng Lin;C. Busso
中科院分区:
其他
文献类型:
--
作者:
Wei-Cheng Lin;C. Busso

文献摘要

相似文献

块级语音情感识别(SER)是一种常用的建模方案,可以获得比句子级公式更好的识别性能。一个关键的开放性问题是词汇边界信息在将句子分成小块的过程中所起的作用。提供精确的词汇边界信息以将语音分割成块(例如,词级对齐)是否有任何好处?本研究通过探索不同的词块级语义切分策略,分析了词汇边界信息在词块级语义切分中的作用。我们比较了六种块级分割策略,这些策略要么考虑词级对齐,要么通过改变块的数量和块的持续时间来考虑传统的基于时间的分割方法。我们进行了大量的实验来评估这些使用多个语料库和多个声学特征集的块级分割方法。结果显示单词级时间边界的贡献很小,其中将块集中在单词周围不会导致显着的性能提升。相反,有效地将句子分割成数据块的关键因素是根据句子中口语单词的数量来定义数据块的数量。
Chunk-level speech emotion recognition (SER) is a common modeling scheme to obtain better recognition performance than sentence-level formulations. A key open question is the role of lexical boundary information in the process of splitting a sentence into small chunks. Is there any benefit in providing precise lexical boundary information to segment the speech into chunks (e.g., word-level alignments)? This study analyzes the role of lexical boundary information by exploring alternative segmentation strategies for chunk-level SER. We compare six chunk-level segmentation strategies that either consider word-level alignments or traditional time-based segmentation methods by varying the number of chunks and the duration of the chunks. We conduct extensive experiments to evaluate these chunk-level segmentation approaches using multiples corpora, and multiple acoustic feature sets. The results show a minor contribution of the word-level timing boundaries, where centering the chunks around words does not lead to significant performance gains. Instead, the critical factor to effectively segment a sentence into data chunks is to define the number of chunks according to the number of spoken words in the sentence.