课题基金 / 基金详情

CRII: CHS: Predicting When, Why, and How Multiple People Will Disagree when Answering a Visual Question

CRII: CHS: Predicting When, Why, and How Multiple People Will Disagree when Answering a Visual Question
CRII:CHS:预测多人在回答视觉问题时何时、为何以及如何产生分歧
批准号:
1755593
负责人:
Danna Gurari
金额:
$17.49万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-01 至 2021-04-30

项目摘要

项目成果

Danna Gurari的其他基金

相似基金

相关文献

中文摘要
翻译
视觉问答(VQA)系统的目标是使人们能够找到关于任何图像的任何问题的答案。例如,VQA系统可以让盲人解决日常的视觉挑战,比如学习一双袜子是否匹配,或者学习罐头里的食物类型。VQA服务还可以促进智能环境的创建,比如在任何给定时间监控工厂装配线上有多少次品。现有VQA系统的一个限制是,它们没有考虑到一个视觉问题可能会从不同的人那里得到不同的答案。如果VQA系统允许用户预测和解决可能出现的任何答案分歧,那么VQA系统可以节省时间并减少用户的挫败感。盲人和视力正常的人可以更快、更准确地了解人类对视觉世界的不同看法。VQA服务还可以教会人们如何提出视觉问题,从而引出期望的答案多样性。该项目将创建人工智能(AI)模型,可以解释群体智能中固有的答案可能的多样性。具体来说,人工智能模型将被设计用来预测人类回答分歧的时间、原因和方式,这反过来将为人机合作关系提供新的设计。这是具有挑战性的,因为它需要设计框架,同时建模和综合不同的和潜在的冲突的感知图像和语言的许多可能的不一致的原因。为了确保人工智能模型在广泛的应用中得到推广,将使用由盲人和视力正常的人提出的超过100万个视觉问题的现有语料库来创建带注释的数据集,这些数据集将表明何时、为什么以及有多少答案出现分歧。然后将开发方法,直接从视觉问题中自动预测答案的多样性,以及为什么会出现分歧。最后,将设计一个系统来指导视障用户更快地制定视觉问题,以便他们能够收到单一的、明确的人群反应(例如,引导人们更好地用手机相机构建感兴趣的视觉内容)。将对盲人用户进行用户研究,以经验检验新系统的有效性,重点是在现实世界的实时情况下发现基于人的问题。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The goal of a visual question answering (VQA) system is to empower people to find the answer to any question about any image. For example, a VQA system could enable blind people to address daily visual challenges such as learning whether a pair of socks match or learning what type of food is in a can. VQA services could also facilitate the creation of smarter environments, say to monitor how many defective products are on a factory assembly line at any given time. A limitation of existing VQA systems is that they do not account for the fact that a visual question may elicit different answers from different people. VQA systems could save time and reduce user frustration if they empowered users to anticipate and resolve any answer disagreements that may arise. Blind and sighted people could more rapidly and accurately learn about the diversity of human perspectives on the visual world. VQA services also could teach people how to ask visual questions that elicit the desired answer diversity.This project will create artificial intelligence (AI) models that can account for the possible diversity of answers inherent in crowd intelligence. Specifically, AI models will be designed to predict when, why, and how human answer disagreement occurs, which in turn will enable new designs for human-computer partnerships. This is challenging because it necessitates designing frameworks that simultaneously model and synthesize different and potentially conflicting perceptions of images and language for the many possible causes of disagreement. To ensure that the AI models generalize across a broad range of applications, an existing corpus of over one million visual questions asked by blind and sighted people will be used to create annotated datasets that indicate when, why, and how much answer disagreement arises. Methods will then be developed for automatically predicting directly from a visual question how much answer diversity will arise from a crowd, and why disagreement arises when it does. Finally, a system will be designed for guiding visually-impaired users to more quickly formulate visual questions so they can receive a single, unambiguous crowd response (e.g., guide the person to better frame the visual content of interest with a mobile phone camera). User studies with blind users will be conducted to empirically test the efficacy of the new system, with a focus on uncovering human-based issues in real-world, real-time situations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr.2018.00380
发表时间: 2018-02
期刊: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子: --
作者: [D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham]
通讯作者: D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham
DOI: 10.1609/hcomp.v6i1.13341
发表时间: 2018-06
期刊:
影响因子: --
作者: [Chun-Ju Yang;K. Grauman;D. Gurari]
通讯作者: Chun-Ju Yang;K. Grauman;D. Gurari
DOI: 10.1109/wacv.2019.00166
发表时间: 2018-03
期刊: 2019 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子: --
作者: [Yinan Zhao;Brian L. Price;Scott D. Cohen;D. Gurari]
通讯作者: Yinan Zhao;Brian L. Price;Scott D. Cohen;D. Gurari
BrowseWithMe: An Online Clothes Shopping Assistant for People with Visual Impairments
BrowseWithMe:为视障人士提供的在线服装购物助手
DOI: 10.1145/3234695.3236337
发表时间: 2018
期刊: ACM SIGACCESS Conference on Computers and Accessibility
影响因子: --
作者: [Stangl, Abigale J., Kothari, Esha, Jain, Suyog D., Yeh, Tom, Grauman, Kristen, Gurari, Danna]
通讯作者: Gurari, Danna
Collaborative Research: SaTC: CORE: Medium: Novel Algorithms and Tools for Empowering People Who Are Blind to Safeguard Private Visual Content
  • 批准号:
    2126297
  • 项目类别:
    Standard Grant
  • 资助金额:
    $56.77万
  • 财政年份:
    2021
  • 负责人:
    Danna Gurari
  • 依托单位:
Collaborative Research: SaTC: CORE: Medium: Novel Algorithms and Tools for Empowering People Who Are Blind to Safeguard Private Visual Content
  • 批准号:
    2148080
  • 项目类别:
    Standard Grant
  • 资助金额:
    $56.77万
  • 财政年份:
    2021
  • 负责人:
    Danna Gurari
  • 依托单位:
国内基金
海外基金
基于CHS-DRGs和诊疗全流程大数据挖掘的子宫肌瘤手术“主路径+支路径”的复合临床路径模式研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    朱文俊
  • 依托单位:
CHS-DRG模式下ICU老年患者CRE医院感染防控对策研究
3,5-双(2-羟基-4-氟-苯基)-1,2,4-噁二唑-铈配合物@CD-MFO-CHS 脑靶向载药纳米粒的制备及抗 AIS脑保护作用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    15.0万元
  • 批准年份:
    2024
  • 负责人:
    张静夏
  • 依托单位:
威尼斯镰刀菌中几丁质合成关键基因Chs调控菌丝体结构与蛋白消 化特性的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    周治彤
  • 依托单位: