Toward a model for lexical access based on acoustic landmarks and distinctive features

Toward a model for lexical access based on acoustic landmarks and distinctive features
复制标题

DOI:
10.1121/1.1458026
复制
发表时间:
2002-04-01
影响因子:
2.4
通讯作者:
Stevens, KN
Stevens, KN
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Stevens, KN

文献摘要

被引文献

相似文献

本文描述了一种模型,其中声学语音信号经过处理,以片段序列的形式产生语音流的离散表示,每个片段都由一组(或捆绑)二进制独特特征来描述。这些独特的特征指定了语言中使用的音位对比,因此特征值的变化可能会生成新单词。该模型是更通用模型的一部分,该模型从该特征表示中导出单词序列,单词在词典中由特征束序列表示。信号的处理分三个步骤进行:(1) 检测信号特定频率范围内的峰值、谷值和不连续性,从而识别声学标志。界标的类型为称为无发音器官特征(例如,[元音]、[辅音]、[连续音])的独特特征子集提供了证据。 (2)从地标附近的信号导出声学参数,为特定咬合架的动作提供证据,并且通过对这些区域中这些参数的选定属性进行采样来提取声学线索。提取的线​​索的选择取决于地标的类型及其出现的环境。 (3) 将步骤 (2) 中获得的线索结合起来,考虑到上下文,以提供与每个标志相关的“咬合架限制”特征的估计(例如,[嘴唇]、[高]、[鼻])。这些受咬合架约束的特征与(1)中的无咬合架特征相结合,构成了形成模型输出的特征束序列。给出了所使用的提示示例和这种选择的理由,以及当由于增强手势(由说话者招募以使对比度更加突出)或由于 相邻片段的手势重叠。 (C) 2002 年美国声学学会。
This article describes a model in which the acoustic speech signal is processed to yield a discrete representation of the speech stream in terms of a sequence of segments, each of which is described by a set (or bundle) of binary distinctive features. These distinctive features specify the phonemic contrasts that are used in the language, such that a change in the value of a feature can potentially generate a new word. This model is a part of a more general model that derives a word sequence from this feature representation, the words being represented in a lexicon by sequences of feature bundles. The processing of the signal proceeds in three steps: (1) Detection of peaks, valleys, and discontinuities in particular frequency ranges of the signal leads to identification of acoustic landmarks. The type of landmark provides evidence for a subset of distinctive features called articulator-free features (e.g., [vowel], [consonant], [continuant]). (2) Acoustic parameters are derived from the signal near the landmarks to provide evidence for the actions of particular articulators, and acoustic cues are extracted by sampling selected attributes of these parameters in these regions. The selection of cues that are extracted depends on the type of landmark and on the environment in which it occurs. (3) The cues obtained in step (2) are combined, taking context into account, to provide estimates of "articulator-bound" features associated with each landmark (e.g., [Lips], [high], [nasal]). These articulator-bound features, combined with the articulator-free features in (1), constitute the sequence of feature bundles that forms the output of the model, Examples of cues that are used, and justification for this selection, are given, as well as examples of the process of inferring the underlying features for a segment when there is variability in the signal due to enhancement gestures (recruited by a speaker to make a contrast more salient) or due to overlap of gestures from neighboring segments. (C) 2002 Acoustical Society of America.