Quantitative analysis of mathematical documents

Quantitative analysis of mathematical documents
复制标题

DOI:
10.1007/s10032-005-0142-y
复制
发表时间:
2005-09
期刊:
International Journal of Document Analysis and Recognition (IJDAR)
影响因子:
--
通讯作者:
S. Uchida;Akihiro Nomura;Masakazu Suzuki
S. Uchida;Akihiro Nomura;Masakazu Suzuki
中科院分区:
其他
文献类型:
--
作者:
S. Uchida;Akihiro Nomura;Masakazu Suzuki

文献摘要

被引文献

相似文献

数学文献从几个角度进行了分析,为数学和其他科学文献的实际OCR的发展。具体地,使用包含690,000个手动地面真实字符的数学文档的大规模数据库来量化四个视点:(i)字符类别的数量,(ii)异常字符(例如,触摸字符),(iii)字符大小变化,以及(iv)数学表达式的复杂性。这些分析的结果澄清了识别数学文件的困难,然后提出了几个有前途的方向来克服它们。
Mathematical documents are analyzed from several viewpoints for the development of practical OCR for mathematical and other scientific documents. Specifically, four viewpoints are quantified using a large-scale database of mathematical documents, containing 690,000 manually ground-truthed characters: (i) the number of character categories, (ii) abnormal characters (e.g., touching characters), (iii) character size variation, and (iv) the complexity of the mathematical expressions. The result of these analyses clarifies the difficulties of recognizing mathematical documents and then suggests several promising directions to overcome them.