IMPACT: centre of competence in text digitisation

IMPACT: centre of competence in text digitisation
复制标题

IMPACT:文本数字化能力中心

DOI:
--
复制
发表时间:
2011
期刊:
The Hip
影响因子:
--
通讯作者:
A. Conteh
A. Conteh
中科院分区:
--
文献类型:
--
作者:
H. Balk;A. Conteh

文献摘要

被引文献

相似文献

A major focus of recent large scale digitisation initiatives has been historical texts, primarily in the form of out-of-copyright newspapers and books. However, the Optical Character Recognition (OCR) software used to translate the scanned images to machine-readable text does not provide satisfactory results for historical documents. This is due to issues inherent in the material such as warped pages, bleed-through, historical fonts, broken and irregular characters, complex layouts, and spelling variants. In the large scale project Improving Access to Text (IMPACT), a European team of scientists, industry partners and digitisation professionals have been working together to enhance existing and develop new approaches to the extraction of text content from historical documents. The project facilitates a successful collaboration between digitisation professionals, based at institutions digitising millions of historical text documents, and scientists in document analysis, language technologies and OCR. This session will detail the work of IMPACT in the context of real life problems faced in the large scale digitisation programmes of libraries and the legacy that the project will leave to foster further research in advancing the state of the art in extracting textual content from historical documents.