Exploiting Speech Understanding in Intelligent Interfaces
Exploiting Speech Understanding in Intelligent Interfaces
批准号:
06044055
负责人:
WARD Nigel
金额:
$2.69万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for international Scientific Research
财政年份:
1994
资助国家:
日本
项目状态:
已结题
起止时间:
1994 至 1995
中文摘要
我们对口语在人机交互中的应用很感兴趣。灵感来自于这样一个事实,即对于人与人之间的互动,即使没有准确识别对方所说的话,也可以进行有意义的交流——这是由于共享知识和互补的沟通渠道,特别是手势和韵律而成为可能的。我们希望在人机界面中利用这一事实。因此,我们正在做三件事:1。使用简单的语音识别来增强图形用户界面,并与键盘、鼠标和触摸屏等其他输入方式很好地集成。构建能够进行简单对话的系统,主要使用韵律线索。来概述一下我们最近的成功:我们推测,日本人有可能只根据韵律线索,而不参考意义,来决定何时产生许多反向通道话语。我们发现,无论是元音的加长、音量的变化,还是能量水平(用来检测对方什么时候结束说话)本身都不能很好地预测什么时候发出aizuchi。最好的预测者是低音阶。具体来说,在检测到某一区域的音高小于。9倍于当地中位音高,持续150毫秒,在至少600毫秒的语音之后,系统预测200毫秒到300毫秒后的合口音,前提是它在前1秒内没有这样做。我们还基于上述决策规则构建了一个实时系统。一个人工助手将对话引导到一个合适的话题,然后打开系统。打开开关后,傀儡的话语和系统的输出混合在一起,产生了对话的一方。我们发现5个被试中没有一个人意识到他的对话伙伴已经部分自动化了。构建工具和收集数据来帮助完成第一和第二件事。
英文摘要
We are interested in the use of spoken language in human-computer interaction. The inspiration is the fact that, for human-human interaction, meaningful exchanges can take place even without accurate recognition of the words the other is saying --- this being possible due to shared knowledge and complementary communication channels, especially gesture and prosody. We want to exploit this fact for man-machine interfaces.Therefore we are doing three things :1. Using simple speech recognition to augment graphical user interfaces, well integrated with other input modalities : keyboard, mouse, and touch screen.2. Building systems able to engage in simple conversations, using mostly prosodic clues. To sketch out our latest success :We conjectured that it would be possible for Japanese to decide when to produce many back-channel utterances based on prosodic clues alone, without reference to meaning.We found thatneither vowel lengthening, volume changes, nor energy level (to detect when the other finished speaking) were by themselves good predictors of when to produce an aizuchi. The best predictor was a low pitch level.Specifically, upon detection of the end of a region of pitch less than.9 times the local median pitch and continuing for 150ms, coming after at least 600ms of speech, the system predicted an aizuchi 200ms to 300ms later, providing it had not done so within the preceding 1 second.We also built a real-time system based on the above decision rule. A human stooge steered the conversation to a suitable topic and then switched on the system. After swich-on the stooge's utterances and the system's outputs, mixed together, produced one side of the conversation. We found that none of the 5 subjects had realized that his conversation partner had become partially automated.3. Building tools and collecting data to help do 1 and 2.
期刊论文(19)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Tajchman,Gary and Dan,Jurafsky and Eric Folder: "Learning Phonological Rule Probabilities from Speech Corpora with Exploratory Computational Phonology" In Proceedings of ACL95. 9-15 (1995)
Tajchman、Gary 和 Dan、Jurafsky 和 Eric Folder:“通过探索性计算音系学从语音语料库学习音系规则概率”,ACL95 论文集。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Gildea, Daniel and Daniel.Jurafsky: "Learning Bias and Phonological Rules Induction" Computational Linguistics. (1995)
Gildea、Daniel 和 Daniel.Jurafsky:“学习偏差和语音规则归纳”计算语言学。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Nigel, WARD: "Using Prosodic Clucs to Decide When to Produce Back-Channel Utterances" CSLP.
Nigel,WARD:“使用韵律线索来决定何时产生 Back-Channel 话语”CSLP。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Jurafsky, Daniel: "A Probabilistic Model of Lexical and Syntactic Access and Disambiguation" Cognitive Science.
Jurafsky,丹尼尔:“词汇和句法访问和消歧的概率模型”认知科学。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Nigel Ward: "An Approach to Tightly-Coupled Syntactic/Semantic Processing for Speech Understanding" Proceedings of the AAAT Workshop on the Integration of Natural Language and Speech Processing. 50-57 (1994)
Nigel Ward:“用于语音理解的紧耦合句法/语义处理方法”自然语言与语音处理集成 AAAT 研讨会论文集。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 15 条
外国語会話能力養成のための対話的反射訓練システム
-
批准号:12040209
-
项目类别:Grant-in-Aid for Scientific Research on Priority Areas (A)
-
资助金额:$1.34万
-
财政年份:2000
-
负责人:WARD Nigel
-
依托单位:
Non-lexical Sounds : a New Interface Modality for Voice-based Information Delivery Systems
-
批准号:11680412
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.6万
-
财政年份:1999
-
负责人:WARD Nigel
-
依托单位:
実時間音声理解応答を利用した機械操作における指動作訓練支援システムの研究
-
批准号:08750301
-
项目类别:Grant-in-Aid for Encouragement of Young Scientists (A)
-
资助金额:$0.7万
-
财政年份:1996
-
负责人:WARD Nigel
-
依托单位:
海外基金