课题基金 / 基金详情

A scheme for continuous speech recognition in a large context based on the human process of spoken language recognition

A scheme for continuous speech recognition in a large context based on the human process of spoken language recognition
基于人类口语识别过程的大上下文连续语音识别方案
批准号:
03452164
负责人:
FUJISAKI Hiroya
金额:
$4.48万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for General Scientific Research (B)
财政年份:
1991
资助国家:
日本
项目状态:
已结题
起止时间:
1991 至 1992

项目摘要

项目成果

FUJISAKI Hiroya的其他基金

相关文献

中文摘要
翻译
目前大多数的自动语音识别系统无法实现识别性能与人类听众相比,因为他们没有注意到人类的口语识别过程。从这个角度来看,本研究调查人类的过程,并将调查结果纳入一个计划,在一个大的背景下连续语音的自动识别。主要研究结果如下:1.人类口语识别过程的实验研究和模型建立以自然语言为刺激,通过对语音、句法和语义信息的控制,获得了以下关于人类口语识别过程的研究结果。(1)根据实验条件和上下文,语音识别的单位从音素和音节到单词和短语变化很大。(2)更大的单位通常需要更低的准确度来正确识别。(3)的量 ...更多信息 识别给定单元所需的声学信息的大小根据上下文的大小和收听者的先验知识而变化很大。(4)心理词汇提取的准确性和速度随听觉、句法、语义和语篇信息的变化而动态变化,在此基础上,本文构建了一个口语记忆过程的模型.一种口语发音自动识别方案的提出与实现基于上述研究结果和模型,本文提出了一种大语境下连续语音的自动识别方案,其特点是:(1)使用多尺度单元和声学特征表示的准确性;(2)使用韵律特征检测词和短语边界;(3)提取句法、语义、和特殊的信息。系统的主要组成部分已经实现.通过对大上下文连续语音中的音素、音节和单词的识别实验,验证了该方案的有效性和可行性。少
英文摘要
Most of the current systems for automatic speech recognition fail to achieve recognition performance comparable to human listeners, since they are constructed without paying attention to the human processes of spoken language recognition. From this point of view, the present study investigates the human processes and incorporates the findings into a scheme for automatic recognition of continuous speech in a large context. The followings are the main results:1. Experimental investigation and modeling of the human processes of spoken language recognitionUsing as stimuli natural utterances with controlled acoustic, syntactic and semantic information, the following findings were obtained on the human processes of spoken language recognition.(1) The unit of speech recognition varies widely from phones and syllables to words and phrases depending on the experimental condition and context.(2) Larger units generally require less accuracy of representation for correct recognition.(3) The amount … More of acoustic information necessary for recognition of a given unit varies widely depending on the size of context and prior knowledge on the part of the listener.(4) The accuracy and speed of access to mental lexicon varies dynamically depending on the acoustic, syntactic, semantic and discourse information available to the listener.Based on these findings, a model has been constructed for the human processes of spoken language recognition.2. Proposal and implementation of a scheme for automatic recognition of spoken language recognitionBased upon the above findings and the model, a scheme for automatic recognition of continuous speech in a large context has been proposed, featuring (1) use of multiple size units and accuracy of acoustic feature representation, (2) use of prosodic features for word and phrase boundary detection, (3) extraction of syntactic, sematic, and idiosyncratic information from a large context. The main components of the system have been implemented.3. Demonstration of the validity of the proposed schemeThe proposed scheme has been tested by recognition experiments of phones, syllables and words in continuous speech with a large context, and the results have confirmed the essential validity and feasibility of the proposed scheme. Less
期刊论文(38)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Sumio Ohno: "Utilization of lexical information at multiple levels in template matching of words and phrases in continuous speech" Reports of 1992 Spring Meeting of the Acoustical Society of Japan. vol. 1. 95-96 (1992)
Sumio Ohno:“在连续语音中单词和短语的模板匹配中多层次词汇信息的利用”日本声学学会 1992 年春季会议的报告。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
峯松,信明: "連続音声知覚における高次言語情報の及ぼす影響" 日本音響学会聴覚研究会資料. H-92-56. 1-6 (1992)
Minematsu, Nobuaki:“高阶语言信息对连续语音感知的影响”日本声学学会听觉研究小组的材料 H-92-56 (1992)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
峯松,信明: "複数の時間的単位・精度の音響的特徴表現を用いた音声認識" 日本音響学会平成4年春季研究発表会講演論文集. 1. 31-32 (1992)
Minematsu、Nobuaki:“使用具有多个时间单位和精度的声学特征表示的语音识别”日本声学学会 1992 年春季会议记录(1992 年)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
共 17 条
    Automatic Estimation of Fundamental Frequency Contour Parameters and Automatic Acquisition of Generative rules
    • 批准号:
      11480090
    • 项目类别:
      Grant-in-Aid for Scientific Research (B).
    • 资助金额:
      $7.81万
    • 财政年份:
      1999
    • 负责人:
      FUJISAKI Hiroya
    • 依托单位:
    Construction of an Intelligent System for information Retrieval in an Environment of Information Network
    • 批准号:
      09558041
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $5.76万
    • 财政年份:
      1998
    • 负责人:
      FUJISAKI Hiroya
    • 依托单位:
    A System for Rule Synthesis of Prosodic Features of Speech of Multiple Language Based on a Generative Model of Fundamental Frequency Contours
    • 批准号:
      08458090
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $4.42万
    • 财政年份:
      1996
    • 负责人:
      FUJISAKI Hiroya
    • 依托单位:
    International Coordination of Speech Databases, Prosodic Labeling, and Speech Input/Output Systems Assessment
    • 批准号:
      08044173
    • 项目类别:
      Grant-in-Aid for international Scientific Research
    • 资助金额:
      $7.36万
    • 财政年份:
      1996
    • 负责人:
      FUJISAKI Hiroya
    • 依托单位: