课题基金 / 基金详情

Camera-based Text Recognition from Complex Backgrounds for the Blind or Visually

Camera-based Text Recognition from Complex Backgrounds for the Blind or Visually
盲人或视觉复杂背景下基于摄像头的文本识别
批准号:
7977496
负责人:
YingLi Tian
金额:
$19.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-09-01 至 2012-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
今天,美国有1000多万盲人和视障人士。计算机视觉、数码相机和便携式计算机的最新技术发展使得通过开发基于相机的 将计算机视觉技术与其他现有产品相结合的产品。尽管一些阅读助手是专门为盲人或视障人士设计的,但阅读复杂背景或非平面的文本 表面问题是一个非常具有挑战性的问题,尚未得到成功解决。许多日常任务都涉及到这些具有挑战性的条件,比如自动售货机上的阅读说明,书架上的书名,药瓶上的说明或汤罐上的标签。 本提案的重点是开发新的计算机视觉算法,以识别来自复杂背景的文本:1)来自具有多种不同颜色的背景(例如,书名排列在书架上)和2)从非平坦的表面(例如药瓶或汤罐上的标签)。新开发的计算机视觉技术将与现成的光学字符识别(OCR)和语音合成软件产品集成在一起。视觉信息将通过头盔捕获 语音显示将通过微型扬声器、耳机或蓝牙设备输出,并由便携式计算机(PDA或手机)分析。将制作一个实用的阅读系统原型,用于阅读复杂背景和非平坦表面的文本。该系统将具有成本效益,因为它只需要一个头戴式摄像头(100万分辨率为100美元),一台可穿戴计算机(300美元),以及两个迷你扬声器或耳机。“ReadIRlS”[74]OCR软件的价格不到150美元,“TextAloud”语音合成软件的价格约为30美元[75]。 该项目将在纽约城市学院(CCNY)和纽约灯塔国际公司执行,为期两年。CCNY位于纽约市哈莱姆区,被指定为少数族裔机构和拉美裔服务机构(37%是西班牙裔,27%是非裔美国人)。国际灯塔是一家领先的非营利性组织,致力于保护视力,提供急需的视力和康复服务,帮助所有年龄段的人克服视力丧失的挑战。在这两年中,我们将1)开发新的算法来识别来自多种不同颜色背景的文本;2)开发新的算法来识别非平坦表面的文本;以及3)通过与现有的光学字符识别(OCR)和语音合成软件产品相集成,为盲人用户开发具有成本效益的原型阅读系统。原型和算法的有效性将由视力正常的人和视力障碍的人来评估。将创建复杂背景(多种颜色和非平坦表面)上的文本数据库,用于算法和系统评估。该数据库将向计算机视觉和视力康复科学领域的研究界提供。总之,这项工作将为下一代盲人阅读助手的设计提供一个基于研究的基础,并产生一个实用的原型,帮助盲人用户在现实环境中阅读复杂背景下的文本。
英文摘要
There are more than 10 million blind and visually impaired people living in America today. Recent technology developments in computer vision, digital cameras, and portable computers make it possible to assist these individuals by developing camera-based products that combine computer vision technology with other existing products. Although a number of reading assistants have been designed specifically for people who are blind or visually impaired, reading text from complex backgrounds or non-flat surfaces is very challenging and has not yet been successfully addressed. Many everyday tasks involve these challenging conditions, such as reading instructions on vending machines, titles of books aligned on a shelf, instructions on medicine bottles or labels on soup cans. This proposal focuses on the development of new computer vision algorithms to recognize text from complex backgrounds: 1) from backgrounds with multiple different colors (e.g .. the titles of books lined up on a shelf) and 2) from non-flat surfaces (e.g .. labels on medicine bottles or soup cans). The newly developed computer vision techniques will be integrated with off-the-shelf optical character recognition (OCR) and speech-synthesis software products. Visual information will be captured via a head-mounted camera (on sunglasses or hat) and analyzed by a portable computer (PDA or cell phone), while the speech display will be outputted via mini speakers, earphones, or Bluetooth device. A practical reading system prototype will be produced to read text from complex backgrounds and non-flat surfaces. The system will be cost-effective since it requires only a head mounted camera (<US$100 for 1M resolution), a wearable computer (<US$300), and two mini-speakers or earphones. The price of "ReadIRlS" [74] OCR software is under $150 and the "TextAloud" speech synthesis software is about $30 [75]. This project will be executed over two years at the City College of New York (CCNY) and Lighthouse International, New York. CCNY, located in the Harlem neighborhood of New York City, is designated as both a Minority Institution and a Hispanic-serving Institution (37% Hispanic and 27% African American). Lighthouse International is a leading non-profit organization dedicated to preserving vision and to providing critically needed vision and rehabilitation services to help people of all ages overcome the challenges of vision loss. During the two years, we will 1) develop new algorithms to recognize text from backgrounds with multiple different colors; 2) develop new algorithms to recognize text from non-flat surfaces; and 3) develop a cost-effective prototype reading system for blind users by integrating with off-the-shelf optical character recognition (OCR) and speech-synthesis software products. The effectiveness of the prototype and algorithms will be evaluated by people with normal vision and people with vision impairment. A database of text on complex backgrounds (multiple colors and non-flat surfaces) will be created for algorithm and system evaluation. The database will be made available to research communities in the areas of computer vision and vision rehabilitation science. In summary, this effort will provide a research-based foundation to inform the design of next generation reading assistants for blind persons, as well as produce a practical prototype to help the blind user read text from complex backgrounds in real-world environments.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/wocc.2011.5872294
发表时间: 2011
期刊: WOCC ... : Wireless & Optical Communications Conference : the ... Annual Wireless & Optical Communications Conference. Annual Wireless & Optical Communications Conference
影响因子: --
作者: [Hasanuzzaman FM, Yang X, Tian Y]
通讯作者: Tian Y
DOI: 10.1007/s00138-012-0431-7
发表时间: 2013-04-01
期刊: MACHINE VISION AND APPLICATIONS
影响因子: 3.3
作者: [Tian, YingLi, Yang, Xiaodong, Yi, Chucai, Arditi, Aries]
通讯作者: Arditi, Aries
DOI: 10.1007/s13721-013-0026-x
发表时间: 2013-07-01
期刊: NETWORK MODELING AND ANALYSIS IN HEALTH INFORMATICS AND BIOINFORMATICS
影响因子: 2.3
作者: [Yi, Chucai, Flores, Roberto W, Chincha, Ricardo, Tian, Yingli]
通讯作者: Tian, Yingli
Detecting Signage and Doors for Blind Navigation and Wayfinding.
检测标牌和门以进行盲导航和寻路。
DOI: 10.1007/s13721-013-0027-9
发表时间: 2013
期刊: Network modeling and analysis in health informatics and bioinformatics
影响因子: 2.3
作者: [Wang,Shuihua, Yang,Xiaodong, Tian,Yingli]
通讯作者: Tian,Yingli
共 11 条
    海外基金