课题基金 / 基金详情

Universal Text Recognition: A Wayfinding Tool for People with Visual Impairments

Universal Text Recognition: A Wayfinding Tool for People with Visual Impairments
通用文本识别:为视觉障碍人士提供的寻路工具
批准号:
7489903
负责人:
Erik G Learned-Miller
金额:
$18.87万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2010-08-31

项目摘要

项目成果

Erik G Learned-Miller的其他基金

相似基金

相关文献

中文摘要
翻译
描述(申请人提供):使用传统的辅助设备(导狗器和长拐杖)和更新的技术,如全球定位系统、印刷文本阅读设备(例如,库兹韦尔-全国盲人文本到语音阅读器联合会)和其他技术的组合,视障人士已经获得了令人印象深刻的自主性。尽管如此,阅读街道标志、商店前横幅、字幕和世界上无处不在的其他形式的文本的愿望或需要,如果没有其他人的帮助,是无法满足的。我们的目标是开发在复杂的室内和室外环境中阅读文本的软件。我们认为,这是通用阅读器中缺少的关键技术,这是一种可以帮助盲人和视障人士在自然场景和环境中导航和操作的设备。我们建议为这样的设备开发新的算法和软件。具体地说,该软件将在高度多样化和复杂的环境中读取数码相机输入的文本,例如街头场景或商业建筑内的文本,并将文本转换为语音或盲文,供盲人使用。我们关注三个核心问题:准确性。最根本的任务是提高基本文本检测和识别算法的准确率。目前的系统根本不能识别足够多的单词,不能用于实际应用。结合用户输入和目标。我们的目标是开发软件机制,使用户可以向设备提供必要的输入,以便在适当的时候适当地缩小图像分析的范围。例如,如果用户可以指定他或她正在寻找“咖啡”,则搜索可以根据用户的请求进行定制,从而降低对与用户目标匹配的文本的检测阈值,并删除不相关的文本。优雅的失败。与提高准确性一样重要的是提供我们所说的优雅故障,即最大限度地减少由于设备错误造成的有害影响,这对此类设备的用户可能更重要。这是此类系统的一个关键且经常被忽视的方面。重要的是,这样的设备的用户不会被误导,使其在设备不正确时(例如,在过马路时)相信它是正确的。由于在户外环境中阅读文本的技术任务可能会任意困难,该设备不可避免地会出现错误。它的主要目标是设计软件,以产生关于返回结果的置信度的反馈,以及将减轻错误影响的其他提示。然后,用户可以评估设备提供的信息的可靠性,并根据当前情况的具体情况做出是否接受结果的明智决定。目前,视力受损的人必须严重依赖于其他有视力的人旅行和日常生活中重要的目的地。该项目的目标是为一种设备生产软件,该设备可以阅读(和说出)标牌、标语牌、字幕和商店正面的单词,供视障用户使用。这样的装置将极大地提高这些人的独立性和自主性。
英文摘要
DESCRIPTION (provided by applicant): Visually impaired individuals have achieved impressive autonomy using a combination of traditional aids (dog guides and long canes) and more recent advances such as global positioning systems, reading devices for printed text (e.g., the Kurzweil-National Federation for the Blind text-to-speech reader), and other technologies. Still, the desire or need to read street signs, store front banners, marquees, and other forms of text that are ubiquitous in the world cannot be met without help from another person. Our goal is to develop software for reading text in complex indoor and outdoor environments. We believe this is the key missing piece of technology in a universal reader, a device that could assist those who are blind and visually impaired in navigating and operating in natural scenes and environments. We propose the development of new algorithms and software for such a device. Specifically, the software will read text from digital camera input in highly diverse and complex environments, such as those found in street scenes or inside commercial buildings and convert that text to speech or Braille for use by people who are blind. We focus on three central issues: Accuracy. The most fundamental task is to increase the accuracy of the basic text detection and recognition algorithms. Current systems simply do not recognize enough words to be practically useful. Incorporating user input and goals. We aim to develop software mechanisms whereby a user can provide essential input to the device to appropriately narrow the image analysis when appropriate. For example, if the user can specify that he or she is seeking "coffee", then the search can be tailored to the user's request, lowering the detection threshold for text that matches the user's goal, and pruning out irrelevant text. Graceful failure. Just as important as increasing accuracy-perhaps even more important to the user of such a device-is to provide what we refer to as graceful failure, i.e. the minimization of harmful effects due to errors made by the device. This is a critical and often overlooked aspect of such systems. It is essential that the user of such a device not be misled into believing that the device is correct when it is not, for example, when crossing the street. Because the technical task of reading text in outdoor environments can be arbitrarily difficult, the device will inevitably make errors. It is a primary goal to design software that produces feedback about the confidence level of returned results, and other cues that will mitigate the impact of errors. The user can then assess the reliability of the information provided by the device and make an intelligent decision about whether to accept the results, depending upon the specifics of the current situation. Currently people who are visually impaired must rely heavily on others who are sighted to travel and destinations that are important for everyday living. The goal of this project is to produce software for a device that can read (and speak) words on signs, placards, marquees, and store fronts to visually impaired users. Such a device would dramatically increase the independence and autonomy of such individuals.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Scene text recognition using similarity and a lexicon with sparse belief propagation.
使用相似性和具有稀疏置信传播的词典进行场景文本识别。
DOI: 10.1109/tpami.2009.38
发表时间: 2009
期刊: IEEE transactions on pattern analysis and machine intelligence
影响因子: 23.6
作者: [Weinman,JerodJ, Learned-Miller,Erik, Hanson,AllenR]
通讯作者: Hanson,AllenR
Universal Text Recognition: A Wayfinding Tool for People with Visual Impairments
海外基金