课题基金 / 基金详情

CRII: CHS: Predicting When, Why, and How Multiple People Will Disagree when Answering a Visual Question

CRII: CHS: Predicting When, Why, and How Multiple People Will Disagree when Answering a Visual Question
CRII:CHS:预测多人在回答视觉问题时何时、为何以及如何产生分歧
批准号:
1755593
负责人:
Danna Gurari
金额:
$17.49万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-01 至 2021-04-30

项目摘要

项目成果

Danna Gurari的其他基金

相似基金

相关文献

中文摘要
翻译
视觉问答 (VQA) 系统的目标是让人们能够找到有关任何图像的任何问题的答案。 例如,VQA 系统可以帮助盲人解决日常视觉挑战,例如了解一双袜子是否匹配或了解罐头中的食物类型。 VQA 服务还可以促进创建更智能的环境,例如监控工厂装配线上在任何给定时间有多少有缺陷的产品。 现有 VQA 系统的局限性在于,它们没有考虑到视觉问题可能会从不同的人那里引出不同的答案。 如果 VQA 系统使用户能够预测并解决可能出现的任何答案分歧,则可以节省时间并减少用户的挫败感。 盲人和视力正常的人可以更快、更准确地了解人类对视觉世界的看法的多样性。 VQA 服务还可以教人们如何提出视觉问题,从而引出所需的答案多样性。该项目将创建人工智能 (AI) 模型,该模型可以解释群体智能中固有的可能答案的多样性。 具体来说,人工智能模型将被设计为预测人类答案分歧何时、为何以及如何发生,这反过来又将为人机合作伙伴关系带来新的设计。 这是具有挑战性的,因为它需要设计一个框架,同时建模和综合不同的、潜在冲突的图像和语言感知,以解决许多可能的分歧原因。 为了确保人工智能模型能够在广泛的应用中推广,将使用由盲人和视力正常的人提出的超过一百万个视觉问题的现有语料库来创建带注释的数据集,以表明出现答案分歧的时间、原因和程度。 然后将开发方法来直接从视觉问题自动预测人群中会产生多少答案多样性,以及为什么会出现分歧。 最后,将设计一个系统来引导视障用户更快地提出视觉问题,以便他们能够收到单一、明确的人群响应(例如,引导人们更好地使用手机摄像头构建感兴趣的视觉内容)。 将针对盲人用户进行用户研究,以实证测试新系统的有效性,重点是在现实世界、实时情况下揭示人为问题。该奖项反映了 NSF 的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The goal of a visual question answering (VQA) system is to empower people to find the answer to any question about any image. For example, a VQA system could enable blind people to address daily visual challenges such as learning whether a pair of socks match or learning what type of food is in a can. VQA services could also facilitate the creation of smarter environments, say to monitor how many defective products are on a factory assembly line at any given time. A limitation of existing VQA systems is that they do not account for the fact that a visual question may elicit different answers from different people. VQA systems could save time and reduce user frustration if they empowered users to anticipate and resolve any answer disagreements that may arise. Blind and sighted people could more rapidly and accurately learn about the diversity of human perspectives on the visual world. VQA services also could teach people how to ask visual questions that elicit the desired answer diversity.This project will create artificial intelligence (AI) models that can account for the possible diversity of answers inherent in crowd intelligence. Specifically, AI models will be designed to predict when, why, and how human answer disagreement occurs, which in turn will enable new designs for human-computer partnerships. This is challenging because it necessitates designing frameworks that simultaneously model and synthesize different and potentially conflicting perceptions of images and language for the many possible causes of disagreement. To ensure that the AI models generalize across a broad range of applications, an existing corpus of over one million visual questions asked by blind and sighted people will be used to create annotated datasets that indicate when, why, and how much answer disagreement arises. Methods will then be developed for automatically predicting directly from a visual question how much answer diversity will arise from a crowd, and why disagreement arises when it does. Finally, a system will be designed for guiding visually-impaired users to more quickly formulate visual questions so they can receive a single, unambiguous crowd response (e.g., guide the person to better frame the visual content of interest with a mobile phone camera). User studies with blind users will be conducted to empirically test the efficacy of the new system, with a focus on uncovering human-based issues in real-world, real-time situations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr.2018.00380
发表时间: 2018-02
期刊: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子: --
作者: [D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham]
通讯作者: D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham
DOI: 10.1609/hcomp.v6i1.13341
发表时间: 2018-06
期刊:
影响因子: --
作者: [Chun-Ju Yang;K. Grauman;D. Gurari]
通讯作者: Chun-Ju Yang;K. Grauman;D. Gurari
DOI: 10.1109/wacv.2019.00166
发表时间: 2018-03
期刊: 2019 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子: --
作者: [Yinan Zhao;Brian L. Price;Scott D. Cohen;D. Gurari]
通讯作者: Yinan Zhao;Brian L. Price;Scott D. Cohen;D. Gurari
BrowseWithMe: An Online Clothes Shopping Assistant for People with Visual Impairments
BrowseWithMe:为视障人士提供的在线服装购物助手
DOI: 10.1145/3234695.3236337
发表时间: 2018
期刊: ACM SIGACCESS Conference on Computers and Accessibility
影响因子: --
作者: [Stangl, Abigale J., Kothari, Esha, Jain, Suyog D., Yeh, Tom, Grauman, Kristen, Gurari, Danna]
通讯作者: Gurari, Danna
Collaborative Research: SaTC: CORE: Medium: Novel Algorithms and Tools for Empowering People Who Are Blind to Safeguard Private Visual Content
  • 批准号:
    2126297
  • 项目类别:
    Standard Grant
  • 资助金额:
    $56.77万
  • 财政年份:
    2021
  • 负责人:
    Danna Gurari
  • 依托单位:
Collaborative Research: SaTC: CORE: Medium: Novel Algorithms and Tools for Empowering People Who Are Blind to Safeguard Private Visual Content
  • 批准号:
    2148080
  • 项目类别:
    Standard Grant
  • 资助金额:
    $56.77万
  • 财政年份:
    2021
  • 负责人:
    Danna Gurari
  • 依托单位:
国内基金
海外基金
基于CHS-DRGs和诊疗全流程大数据挖掘的子宫肌瘤手术“主路径+支路径”的复合临床路径模式研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    朱文俊
  • 依托单位:
CHS-DRG模式下ICU老年患者CRE医院感染防控对策研究
3,5-双(2-羟基-4-氟-苯基)-1,2,4-噁二唑-铈配合物@CD-MFO-CHS 脑靶向载药纳米粒的制备及抗 AIS脑保护作用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    15.0万元
  • 批准年份:
    2024
  • 负责人:
    张静夏
  • 依托单位:
威尼斯镰刀菌中几丁质合成关键基因Chs调控菌丝体结构与蛋白消 化特性的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    周治彤
  • 依托单位: