Universal Text Recognition: A Wayfinding Tool for People with Visual Impairments
Universal Text Recognition: A Wayfinding Tool for People with Visual Impairments
批准号:
7489903
负责人:
Erik G Learned-Miller
金额:
$18.87万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2010-08-31
关键词:
AddressAlgorithmsBudgetsCanesCanis familiarisCoffeeCommunications aid for the blindComplexComputer softwareCuesDestinationsDetectionDevelopmentDevicesEngineeringEnvironmentEyeFacility Construction Funding CategoryFailureFeedbackGoalsGrantHousingImageImage AnalysisIndividualLearningLifePersonsPositioning AttributePrincipal InvestigatorPrintingReaderReadingResearchRunningSoftware DesignSpecialistSpecific qualifier valueSpeechSystemTechnologyTextTimeTravelVisionVisual impairmentVisually Impaired PersonsWorkblindbrailledesiredigitalprogramssoftware developmenttoolway finding
中文摘要
描述(由申请人提供):视力受损的个人已经实现了令人印象深刻的自主性,使用传统的辅助工具(狗指南和长手杖)和更新的进步,如全球定位系统,印刷文本的阅读设备(例如,Kurzweil-国家盲人文本到语音阅读器联合会)以及其他技术。尽管如此,如果没有另一个人的帮助,阅读世界上无处不在的街道标志、商店前横幅、大帐篷和其他形式的文本的愿望或需要是无法满足的。我们的目标是开发在复杂的室内和室外环境中阅读文本的软件。我们相信这是通用阅读器中缺少的关键技术,该设备可以帮助盲人和视障人士在自然场景和环境中导航和操作。我们建议开发新的算法和软件,这样的设备。具体来说,该软件将在高度多样化和复杂的环境中从数码相机输入中读取文本,例如在街道场景或商业建筑内发现的文本,并将文本转换为语音或盲文供盲人使用。我们专注于三个核心问题:准确性。最根本的任务是提高基本文本检测和识别算法的准确性。目前的系统只是不能识别足够的单词,实际上是有用的。整合用户输入和目标。我们的目标是开发软件机制,使用户可以向设备提供必要的输入,以便在适当的时候适当地缩小图像分析范围。例如,如果用户可以指定他或她正在寻找“咖啡”,则可以针对用户的请求定制搜索,降低匹配用户目标的文本的检测阈值,并修剪掉不相关的文本。优雅的失败。与提高准确性同样重要的是--也许对这种设备的用户更重要的是--提供我们所说的“优雅故障”,即最小化由于设备所产生的错误而造成的有害影响。这是这种系统的一个关键而又常常被忽视的方面。至关重要的是,这种设备的用户不被误导而相信该设备是正确的,而实际上它不是正确的,例如,当过马路时。因为在户外环境中阅读文本的技术任务可能是任意困难的,所以设备将不可避免地犯错误。设计软件的主要目标是产生关于返回结果的置信度的反馈,以及其他将减轻错误影响的线索。然后,用户可以评估设备提供的信息的可靠性,并根据当前情况的具体情况,做出是否接受结果的明智决定。目前,视力受损的人必须严重依赖其他视力正常的人来旅行,去对日常生活很重要的目的地。这个项目的目标是为一个设备制作软件,该设备可以为视障用户阅读(和说出)标志,标语牌,大帐篷和店面上的文字。这种装置将大大提高这些人的独立性和自主性。
英文摘要
DESCRIPTION (provided by applicant): Visually impaired individuals have achieved impressive autonomy using a combination of traditional aids (dog guides and long canes) and more recent advances such as global positioning systems, reading devices for printed text (e.g., the Kurzweil-National Federation for the Blind text-to-speech reader), and other technologies. Still, the desire or need to read street signs, store front banners, marquees, and other forms of text that are ubiquitous in the world cannot be met without help from another person. Our goal is to develop software for reading text in complex indoor and outdoor environments. We believe this is the key missing piece of technology in a universal reader, a device that could assist those who are blind and visually impaired in navigating and operating in natural scenes and environments. We propose the development of new algorithms and software for such a device. Specifically, the software will read text from digital camera input in highly diverse and complex environments, such as those found in street scenes or inside commercial buildings and convert that text to speech or Braille for use by people who are blind. We focus on three central issues: Accuracy. The most fundamental task is to increase the accuracy of the basic text detection and recognition algorithms. Current systems simply do not recognize enough words to be practically useful. Incorporating user input and goals. We aim to develop software mechanisms whereby a user can provide essential input to the device to appropriately narrow the image analysis when appropriate. For example, if the user can specify that he or she is seeking "coffee", then the search can be tailored to the user's request, lowering the detection threshold for text that matches the user's goal, and pruning out irrelevant text. Graceful failure. Just as important as increasing accuracy-perhaps even more important to the user of such a device-is to provide what we refer to as graceful failure, i.e. the minimization of harmful effects due to errors made by the device. This is a critical and often overlooked aspect of such systems. It is essential that the user of such a device not be misled into believing that the device is correct when it is not, for example, when crossing the street. Because the technical task of reading text in outdoor environments can be arbitrarily difficult, the device will inevitably make errors. It is a primary goal to design software that produces feedback about the confidence level of returned results, and other cues that will mitigate the impact of errors. The user can then assess the reliability of the information provided by the device and make an intelligent decision about whether to accept the results, depending upon the specifics of the current situation. Currently people who are visually impaired must rely heavily on others who are sighted to travel and destinations that are important for everyday living. The goal of this project is to produce software for a device that can read (and speak) words on signs, placards, marquees, and store fronts to visually impaired users. Such a device would dramatically increase the independence and autonomy of such individuals.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Scene text recognition using similarity and a lexicon with sparse belief propagation.
使用相似性和具有稀疏置信传播的词典进行场景文本识别。
DOI:
10.1109/tpami.2009.38
发表时间:
2009
期刊:
IEEE transactions on pattern analysis and machine intelligence
影响因子:
23.6
作者:
[Weinman,JerodJ, Learned-Miller,Erik, Hanson,AllenR]
通讯作者:
Hanson,AllenR
Universal Text Recognition: A Wayfinding Tool for People with Visual Impairments
-
批准号:7298443
-
项目类别:
-
资助金额:$22.88万
-
财政年份:2007
-
负责人:Erik G Learned-Miller
-
依托单位:
海外基金