Offline recognition of unconstrained handwritten texts using HMMs and statistical language models

Offline recognition of unconstrained handwritten texts using HMMs and statistical language models
复制标题

DOI:
10.1109/tpami.2004.14
复制
发表时间:
2004-06
影响因子:
23.6
通讯作者:
A. Vinciarelli;Samy Bengio;H. Bunke
A. Vinciarelli;Samy Bengio;H. Bunke
中科院分区:
计算机科学1区
文献类型:
--
作者:
A. Vinciarelli;Samy Bengio;H. Bunke

文献摘要

被引文献

相似文献

本文提出了一种大词汇量无限制手写文本的离线识别系统。关于这些数据的唯一假设是,它是用英语写的。这允许应用统计语言模型,以提高我们系统的性能。已经使用单个和多个写入器数据执行了几个实验。使用了可变大小(从10,000到50,000字)的词汇表。语言模型的使用被证明提高了系统的准确性(当词典包含50,000个单词时,错误率对于单个写入者数据降低了/SPL sim/50%,对于多写入者数据降低了/SPL sim/25%)。详细描述了我们的方法,并与文献中提出的处理相同问题的其他方法进行了比较。提出了一种能够正确处理无约束文本识别的实验方案。
This paper presents a system for the offline recognition of large vocabulary unconstrained handwritten texts. The only assumption made about the data is that it is written in English. This allows the application of statistical language models in order to improve the performance of our system. Several experiments have been performed using both single and multiple writer data. Lexica of variable size (from 10,000 to 50,000 words) have been used. The use of language models is shown to improve the accuracy of the system (when the lexicon contains 50,000 words, the error rate is reduced by /spl sim/50 percent for single writer data and by /spl sim/25 percent for multiple writer data). Our approach is described in detail and compared with other methods presented in the literature to deal with the same problem. An experimental setup to correctly deal with unconstrained text recognition is proposed.