课题基金 / 基金详情

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
项目摘要 在美国,大约有1200万人被诊断为视力障碍。这些 在我们的现代环境中,个人面临着独特的挑战,在这个环境中,许多关键信息与 教育、就业、娱乐和社区以数字视频的形式呈现。无法访问 信息可能导致社会排斥或危及生命,如果个人需要访问它以 做出与他们的健康和安全相关的决定。例如,在个人或全球健康危机中,个人 可能需要访问通过视频或动态信息图表传达的海量信息,以便 做出明智的决定。为了满足这一需求,在线平台YouDescribe允许盲人和低视力 (BLV)用户请求业余志愿者创建视频描述,也称为音频描述 (广告),YouTube视频。然而,该平台已经跟不上压倒性的需求, YouDescribe用户愿望清单上92.5%的视频仍未被描述。这项提案的总体目标是 就是要建立一个人工智能驱动的系统,适合大范围使用,自动生成在线描述 视频,以及回答BLV用户关于视频内容的问题。这样做的理由是 项目是,基于人工智能的工具是必要的,以便于及时访问出现在 每天都在上网。拟议的工作包括三个具体目标:1)开发一个基于AI的工具 与目视描述者合作,更高效地生成视频描述并增加 提供可访问的视频。我们的目标是创建一个人工智能驱动的NarariBot,它将减少 新手志愿者需要80%的视频描述;2)在协作中开发基于AI的工具 与BLV个人合作,提供在线视频中用户驱动的可视信息访问。我们的目标是发展 一个AI驱动的QABot,允许用户暂停视频,询问内容,并立即接收 回答(例如,“狗是什么品种的?”,“德国牧羊犬”),准确率达到80%;以及3)发展 并公开发布大规模数据集,以改善视频可访问性的机器学习。这些小说 将使用数据集来提高NarariBot和QABot的质量和精度,直到人工智能生成 描述和答案只需要人类志愿者的最小干预,可以直接为BLV用户服务。 这项拟议的研究具有创新性,因为它专注于视频,而现有的人工智能驱动的努力旨在解决 这个问题主要集中在静态照片或图像上。这也是为数不多的几项直接 与BLV人员合作开发人工智能驱动的系统,以产生视觉描述或回答视觉 问题。这项拟议的研究意义重大,因为它将产生开源的、人工智能驱动的工具,将提供 BLV个人对其独立导航信息丰富的世界的能力进行了前所未有的控制 在线视频,从而改善他们的健康和福祉。
英文摘要
Project Summary Approximately 12 million people in the United States have been diagnosed with a visual impairment. These individuals face unique challenges in our modern environment, where much critical information related to education, employment, entertainment, and community is presented in the form of digital videos. Inaccessible information can result in social exclusion or become life threatening if individuals require access to it in order to make decisions related to their health and safety. For example, in a personal or global health crisis, individuals may need to access the mass amounts of information conveyed via videos or dynamic infographics in order to make informed decisions. To address this need, the online platform YouDescribe allows blind and low vision (BLV) users to request amateur volunteers to create video descriptions, also referred to as audio descriptions (AD), of YouTube videos. However, the platform has been unable to keep up with the overwhelming demand, and 92.5% of videos on the YouDescribe user wish list remain undescribed. The overall objective of this proposal is to build an AI-driven system, suitable for use on a wide-scale, to automatically generate descriptions of online videos, as well as answer questions asked by BLV users about the content of videos. The rationale for this project is that AI-based tools are necessary to facilitate timely access to the deluge of new videos appearing on the Internet every day. The proposed work encompasses three specific aims: 1) develop an AI-based tool in collaboration with sighted describers that more efficiently produces video descriptions and increases the availability of accessible videos. The goal is to create an AI-driven NarrationBot that will decrease the time required for novice volunteers to produce video descriptions by 80%; 2) develop an AI-based tool in collaboration with BLV individuals that offers user-driven access to visual information in online videos. The goal is to develop an AI-driven QABot that allows users to pause a video, ask questions about content, and receive immediate answers (e.g., “What breed is the dog?”, “German shepherd”) that are accurate 80% of the time; and 3) develop and publicly release large-scale datasets to improve machine learning for video accessibility. These novel datasets will be used to increase the quality and accuracy of NarrationBot and QABot until AI-generated descriptions and answers need minimal intervention from human volunteers and can serve BLV users directly. The proposed research is innovative because it focuses on videos, whereas existing AI-driven efforts to address this problem have focused primarily on static photos or images. It is also one of only a few efforts to directly partner with BLV individuals to develop AI-driven systems that produce visual descriptions or answer visual questions. The proposed research is significant because it will result in open-source, AI-driven tools that will give BLV individuals unprecedented control over their ability to independently navigate the information-rich world of online videos, thus improving their health and wellbeing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金