Cascading and Direct Approaches to Unsupervised Constituency Parsing on Spoken Sentences

Cascading and Direct Approaches to Unsupervised Constituency Parsing on Spoken Sentences
复制标题

口语句子无监督选区解析的级联和直接方法

DOI:
10.1109/icassp49357.2023.10094575
复制
发表时间:
2023
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Hung
Hung
中科院分区:
--
文献类型:
--
作者:
Yuan Tseng;Cheng;Hung

文献摘要

被引文献

相似文献

过去关于无监督解析的工作仅限于书面形式。在本文中,我们首次研究了在未标记的口语句子和未配对的文本数据下的无监督口语选区分析。目标是以选区解析树的形式确定口语句子的分层句法结构,这样每个节点都是与一个成分对应的音频范围。我们比较了两种方法:(1)将无监督自动语音识别(ASR)模型与无监督解析器级联,以获得ASR转录本上的解析树;(2)在连续词级语音表示上直接训练无监督解析器。这是通过首先将话语分割成词级片段序列,并在片段内聚合自监督语音表示以获得片段嵌入来完成的。我们发现,在未配对文本上单独训练解析器,并直接将其应用于ASR转录本进行推理,对于无监督解析可以产生更好的结果。此外,我们的结果表明,准确的分割本身可能足以准确地解析口语句子。最后,我们证明了直接方法可以正确地学习头-起始和头-结束语言的头-方向性,而没有任何明显的归纳偏差。
Past work on unsupervised parsing is constrained to written form. In this paper, we present the first study on unsupervised spoken constituency parsing given unlabeled spoken sentences and unpaired textual data. The goal is to determine the spoken sentences’ hierarchical syntactic structure in the form of constituency parse trees, such that each node is a span of audio that corresponds to a constituent. We compare two approaches: (1) cascading an unsupervised automatic speech recognition (ASR) model and an unsupervised parser to obtain parse trees on ASR transcripts, and (2) direct training an unsupervised parser on continuous word-level speech representations. This is done by first splitting utterances into sequences of word-level segments, and aggregating self-supervised speech representations within segments to obtain segment embeddings. We find that separately training a parser on the unpaired text and directly applying it on ASR transcripts for inference produces better results for unsupervised parsing. Additionally, our results suggest that accurate segmentation alone may be sufficient to parse spoken sentences accurately. Finally, we show the direct approach may learn head-directionality correctly for both head-initial and head-final languages without any explicit inductive bias.