Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition
Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition
复制标题
DOI:
10.1109/icassp49357.2023.10096861
复制
发表时间:
2023-06
期刊:
影响因子:
--
通讯作者:
Wei-Cheng Lin;C. Busso
中科院分区:
文献类型:
--
作者:
Wei-Cheng Lin;C. Busso
Chunk-level speech emotion recognition (SER) is a common modeling scheme to obtain better recognition performance than sentence-level formulations. A key open question is the role of lexical boundary information in the process of splitting a sentence into small chunks. Is there any benefit in providing precise lexical boundary information to segment the speech into chunks (e.g., word-level alignments)? This study analyzes the role of lexical boundary information by exploring alternative segmentation strategies for chunk-level SER. We compare six chunk-level segmentation strategies that either consider word-level alignments or traditional time-based segmentation methods by varying the number of chunks and the duration of the chunks. We conduct extensive experiments to evaluate these chunk-level segmentation approaches using multiples corpora, and multiple acoustic feature sets. The results show a minor contribution of the word-level timing boundaries, where centering the chunks around words does not lead to significant performance gains. Instead, the critical factor to effectively segment a sentence into data chunks is to define the number of chunks according to the number of spoken words in the sentence.