课题基金 / 基金详情

Non-lexical Sounds : a New Interface Modality for Voice-based Information Delivery Systems

Non-lexical Sounds : a New Interface Modality for Voice-based Information Delivery Systems
非词汇声音:基于语音的信息传递系统的新接口模式
批准号:
11680412
负责人:
WARD Nigel
金额:
$1.6万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1999
资助国家:
日本
项目状态:
已结题
起止时间:
1999 至 2000

项目摘要

项目成果

WARD Nigel的其他基金

相似基金

相关文献

中文摘要
翻译
本报告中收集的论文描述了在1999财年-2000财年资助的名为“非词汇声音:基于语音的信息传递系统的新接口模式”的资助下进行的研究。*呃,过去5年,交互式语音应答(IVR)系统在美国变得无处不在,并在日本取得进展。事实上,从真人那里获取火车时刻表、公寓信息、呼叫路线、航班信息、天气信息等变得越来越不可能。这些系统普遍不受欢迎的一个原因是需要收听菜单、在菜单之间导航并按下按钮来选择内容,但由于语音识别技术的部署,这一困难正在得到解决。这些系统的第二个问题是,提供的信息是以固定的块提供的,持续时间从几秒到几十秒,用户基本上被迫在系统回放块时进行收听。相比之下,人们支持…电话上的视频信息越多,灵活性就越大。这其中的一个方面是,他们对听者的反馈做出反应。特别是,在许多对话类型中,听者经常产生反向通道,如嗯,嗯,是,哦,嗯,嗯,等等,信息提供者相应地调整他的陈述。因此,本项目的目的是发现这种非词汇会话声音的意义和功能,并将其用于基于语音的信息传递系统中。本项目的第一部分是对日语和英语两种语言对话中非词汇声音的意义和功能的基础研究。这使人们对这些声音有了新的理解。对于英语,这被正式化为一个模型,在这个模型中,这些项目不是作为固定的单词来解释的,而是作为动态的创造,由一个由10个组成部分的声音和2个组合规则组成的简单模型产生。该模型的语义成分是声音-符号:每个成分的声音都有一定的意义或功能*在咕噜声和语境中是相当恒定的,而会话中咕噜声的意思很大程度上是它的语音成分的意义的总和。项目的第二个成分是在教程系统中使用这些声音。该系统产生非词汇声音,根据上下文的不同而适当地变化,展示了它们在面向目标的对话中的效用。该项目的第三部分是研究非词汇声音在实时控制应用中的使用,其中计算机对来自用户的非词汇建议做出实时响应,例如mm-mm-mm或ack!。仍然留在议程上的主题是:1.非词汇声音的韵律特征的含义模型,2.会话日语中的非词汇声音的模型,3.在目录辅助类型的IVR系统中在给号码期间响应这些声音的系统,以及4.基于这些声音调整其呈现的口语对话Salestalk类型的系统。较少
英文摘要
The papers collected in this report describe research performed under a grant entitled "Non-lexical Sounds : a New Interface Modality for Voice-based Information Delivery Systems" funded for fiscal 1999--2000.*er the past 5 years interactive voice response (IVR) systems have become ubiquitous in the United States and are making inroads in Japan. Indeed, it is becoming impossible to get train schedules, apartment information, call routing, flight information, weather information, and so on from a real person. One reason these systems are universally hated is the need to listen to menus, navigate through them, and push buttons to select content, but this difficulty is being resolved thanks to the deployment of speech recognition technology. The second problem with these systems is that the information provided is given in fixed chunks, lasting from a few seconds to a few tens of seconds, and the user is essentially forced to listen as the system plays back a chunk.In contrast, people pro … More viding information over the telephone are much more flexible. One aspect of this is that they are responsive to feedback from the listener. In particular, in many dialog types the listener frequently produces back-channels, such as uh-huh, uh, yeah-yeah, oh, ummmm and so on, and the information provider adapts his presentation in response.Thus the aim of this project was to discover the meanings and functions of such non-lexical conversational sounds, and to exploit them in voice-based information delivery systems.The first component of the project was basic research into the meanings and functions of non-lexical sounds in conversation in two languages, Japanese and English. This led to a new understanding of these sounds. For English, this was formalized as a model in which these items are explained, not as fixed words, but as dynamic creations, generated by a simple model consisting of 10 component sounds and 2 combining rules. The semantic component of the model is sound-symbolic : each of these component sounds bears some meaning or function * is fairly constant across grunts and across contexts, and the meaning of a conversational grunt is largely the sum of the meanings of its phonetic components.The second component of the project was an use of these sounds in a tutorial system. This system produced non-lexical sounds, suitably varied according to context, demonstrating their utility in goal-oriented dialogs.The third component of the project was a study of the use of non-lexical sounds in a real-time control application, in which the computer responded in real-time to non-lexical advice from the user, such as mm--mm--mm or ack! .Topics which remain on the agenda are : 1. a model of the meanings of the prosodic features of non-lexical sounds, 2. a model of non-lexical sounds in conversational Japanese, 3. a system which responds to these sounds during number-giving in a directory-assistance type IVR system, and 4. a spoken dialog salestalk-type system which adapts its presentation based on these sounds. Less
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
Nigel Ward: "Sound Symbolism in Conversational Grunts in English (submitted)"Language and Speech. (submitted). (2001)
奈杰尔·沃德(Nigel Ward):“英语会话咕噜声中的声音象征(已提交)”语言和言语。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Wataru Tsukahara and Nigel Ward: "Responding to Subtle, Fleeting Changes in the User's Internal State"CHI 2001 Conference on Human Factors in Computing Systems. (to appear).
Wataru Tsukahara 和 Nigel Ward:“响应用户内部状态中微妙、短暂的变化”CHI 2001 年计算系统人为因素会议。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Nigel Ward: "The Challenge of Non-lexical Speech Sounds"International Conference on Spoken Language Processing, 2000. 571-574 (2000)
Nigel Ward:“非词汇语音的挑战”国际口语处理会议,2000. 571-574 (2000)
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Wataru Tsukahara and Nigel Ward: "Evaluating Responsiveness in Spoken Dialog Systems"International Conference of Spoken Language Processing 2000. 1097-1100 (2000)
Wataru Tsukahara 和 Nigel Ward:“评估口语对话系统中的响应性”2000 年口语处理国际会议。1097-1100 (2000)
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
8
    外国語会話能力養成のための対話的反射訓練システム
    • 批准号:
      12040209
    • 项目类别:
      Grant-in-Aid for Scientific Research on Priority Areas (A)
    • 资助金额:
      $1.34万
    • 财政年份:
      2000
    • 负责人:
      WARD Nigel
    • 依托单位:
    実時間音声理解応答を利用した機械操作における指動作訓練支援システムの研究
    • 批准号:
      08750301
    • 项目类别:
      Grant-in-Aid for Encouragement of Young Scientists (A)
    • 资助金额:
      $0.7万
    • 财政年份:
      1996
    • 负责人:
      WARD Nigel
    • 依托单位:
    Exploiting Speech Understanding in Intelligent Interfaces
    • 批准号:
      06044055
    • 项目类别:
      Grant-in-Aid for international Scientific Research
    • 资助金额:
      $2.69万
    • 财政年份:
      1994
    • 负责人:
      WARD Nigel
    • 依托单位:
    海外基金