课题基金 / 基金详情

Eye Gaze in Salience Modeling for Robust Spoken Language Understanding

Eye Gaze in Salience Modeling for Robust Spoken Language Understanding
用于鲁棒口语理解的显着性建模中的眼睛注视
批准号:
0535112
负责人:
Joyce Chai
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-11-15 至 2009-10-31

项目摘要

项目成果

Joyce Chai的其他基金

相似基金

相关文献

中文摘要
翻译
在口语对话系统中,由于有限的语音识别和语言理解性能,解释用户语音输入仍然是一个巨大的挑战。如果用户有口音或在嘈杂的环境中说话,这个问题就会进一步放大。然而,先前的研究表明,在多通道系统中,融合两个或更多信息源可以是减少识别不确定性的有效手段,例如通过相互消除歧义。受早期多通道系统研究的启发,在这个项目中,PI将研究眼睛凝视在人机对话中的作用,特别是在为稳健的口语理解而建立的显著模型中。认知研究表明,人的眼睛凝视是一个人“在想什么”的可靠指标之一。具体地说,眼睛凝视与人类的语言处理密切相关。先前的心理语言学研究表明,几乎在听到一个单词后,眼睛会立即转移到对应的现实世界中的参照物。就在说一句话之前,眼睛也会移到提到的物体上。眼睛凝视不仅是高度可靠的,也是一种隐含的、潜意识的言语反射。用户不需要做出有意识的决定;眼睛自动移动到相关对象,甚至用户都没有意识到。受这些心理语言学发现的启发,PI的假设是,在人机对话中,用户的眼睛注视信息与对话语境相结合,可以发出物理世界的一部分(与领域和图形界面相关)在每个交流点上最突出的信号,因此它可能被用来定制对语音输入的解释。基于这一假设,PI将寻求通过一个新的基于突显的框架来提高对话界面中的口语理解,其目标有两个:(1)更好地理解眼睛凝视在人类语言产生中的作用及其在自动输入解释的突显建模中的含义;(2)开发将基于计算凝视的突显建模应用于稳健的口语理解的算法和系统。这些目标将在以下四个方向实现:(A)通过心理语言学研究,调查人眼凝视的效用及其对人机对话过程中的突显建模的影响;(B)开发将眼睛凝视与对话环境相结合的计算突显模型,以在每个交流点自动识别物理世界的显著部分;(C)开发应用新的突显模型来限制强健口语理解的假设空间的方法;以及(D)评估新方法在两个不同应用程序中的通用性:基于3D渲染界面的室内设计/培训应用程序,以及使用基于2D地图的界面的信息搜索应用程序。这些技术将使各种不同的用户受益,尤其是不能用手与图形界面交互的个人(例如,运动障碍用户)。由于这项工作的一个主要应用领域是电子培训和电子学习,因此拟议研究的教育和外联影响可能是深远的;国际和平研究所将作出具体努力,将研究成果转化为课堂。该项目还将为计算机科学、心理学和认知科学的学生提供一个独特的合作机会,从而将密歇根州立大学的多学科研究活动协同起来。
英文摘要
In spoken dialog systems, interpreting user speech input is still a significant challenge due to limited speech recognition and language understanding performance. This problem is further amplified if a user has an accent or is speaking in a noisy environment. However, previous research has shown that, in multimodal systems, fusing two or more information sources can be an effective means of reducing recognition uncertainties, for example through mutual disambiguation. Inspired by earlier work on multimodal systems, in this project the PI will investigate the role of eye gaze in human machine conversation, in particular in salience modeling for robust spoken language understanding. Cognitive studies have shown that human eye gaze is one of the reliable indicators of what a person is "thinking about." Specifically, eye gaze is tightly linked to human language processing. Previous psycholinguistic work has shown that almost immediately after hearing a word, the eyes move to the corresponding real-world referent. And right before speaking a word, the eyes also move to the mentioned object. Not only is eye gaze highly reliable, it is also an implicit, subconscious reflex of speech. The user does not need to make a conscious decision; the eye automatically moves towards the relevant object, without the user even being aware. Motivated by these psycholinguistic findings, the PI's hypothesis is that during human machine conversation user eye gaze information coupled with conversation context can signal a part of the physical world (related to the domain and the graphical interface) that is most salient at each point of communication, thus it can potentially be used to tailor the interpretation of speech input. Based on this hypothesis, the PI will seek to improve spoken language understanding in conversational interfaces through a new salience-based framework with two objectives: (1) To better understand the role of eye gaze in human language production and its implications in salience modeling for automated input interpretation; and (2) To develop algorithms and systems that apply computational gaze based salience modeling to robust spoken language understanding. These objectives will be pursued in the following four directions: (a) Investigation of the utility of human eye gaze and its implications for salience modeling during human machine conversation through psycholinguistic studies; (b) Development of computational salience models that integrate eye gaze with conversation context to automatically identify a salient part of the physical world at each point of communication; (c) Development of approaches that apply the new salience models to constrain the hypothesis space for robust spoken language understanding; and (d) Evaluation of the generality of the new approaches in two different applications: an interior design/training application based on a 3D rendered interface, and an information seeking application using a 2D map-based interface.Broader Impacts: The technologies to be developed in this interdisciplinary project can be applied to many applications such as virtual training systems where users can see the interface and talk to the computer system at the same time. The technologies will benefit a variety of diverse users, and particularly individuals who are unable to interact with graphical interfaces with their hands (e.g., motion disabled users). Since one major application area of the work is e-training and e-learning, the education and outreach impact of the proposed research is potentially profound; the PI will make specific efforts to transfer the research results into classrooms. The project will also provide a unique opportunity for students in Computer Science, Psychology, and Cognitive Science to work together, and thus will synergize multidisciplinary research activities at Michigan State University.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NRI: INT: COLLAB: Collaborative Task Planning and Learning through Language Communication in a Human-Robot Team
NRI: INT: COLLAB: Collaborative Task Planning and Learning through Language Communication in a Human-Robot Team
  • 批准号:
    1830244
  • 项目类别:
    Standard Grant
  • 资助金额:
    $76.84万
  • 财政年份:
    2018
  • 负责人:
    Joyce Chai
  • 依托单位:
RI: Small: Extending Verb Semantics with Causality towards Physical World
  • 批准号:
    1617682
  • 项目类别:
    Standard Grant
  • 资助金额:
    $48.54万
  • 财政年份:
    2016
  • 负责人:
    Joyce Chai
  • 依托单位:
WORKSHOP: Student Consortium at the 2014 ACM Conference on Intelligent User Interfaces
  • 批准号:
    1415879
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.93万
  • 财政年份:
    2013
  • 负责人:
    Joyce Chai
  • 依托单位:
海外基金