Exploring Learning Approaches for Ancient Greek Character Recognition with Citizen Science Data

Exploring Learning Approaches for Ancient Greek Character Recognition with Citizen Science Data
复制标题

利用公民科学数据探索古希腊字符识别的学习方法

DOI:
10.1109/escience51609.2021.00023
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Swindall M
Swindall M
中科院分区:
--
文献类型:
--
作者:
Swindall M

文献摘要

参考文献

被引文献

相似文献

手写字符识别的中心法则仍然与印刷媒体的光学字符识别方法密不可分。除了依赖专有数据和缺乏开放获取软件之外,这些光学字符识别方法对低质量文档(例如,被破坏的)仍然是未知的。在本文中,我们比较和对比了用于打印的最先进的光学字符识别工具的性能,以及与手写输入训练的最先进的机器学习工具包相结合的学习模型。使用Tesseract OCR作为基线,我们构建、优化和评估了三种类型的卷积神经网络,这些网络是在AL-ALL和AL-PUB数据集上训练的,AL-PUB数据集是由志愿者通过Ancient Lives在线公民科学项目标记的手写古希腊字符图像的集合。我们发现我们的机器学习模型的准确率为92.57%,而Tesseract OCR的准确率为11.15%。在我们的分析之后,我们对我们的模型的缺点进行了简要的检查,介绍了公开可用的AL-PUB数据集,并描述了Theia,一个基于Web的工具,它使我们的机器学习模型民主化,供公众使用。最后,我们讨论了我们的研究结果对推进机器学习,手稿转录和数字人文学科交叉研究的承诺。
The central dogma of handwritten character recognition remains inextricably linked to optical character recognition methods for print media. Alongside their reliance on proprietary data and lack of open-access software, the applicability of these optical character recognition methods to handwritten characters from low-quality documents (e.g., that are damaged) remains unknown. In this paper, we compare and contrast the performance of state-of-the-art optical character recognition tools for print and learning models engineered with state-of-the-art machine learning toolkits trained on handwritten inputs. Using Tesseract OCR as a baseline, we build, optimize, and evaluate three types of convolutional neural networks that are trained on the AL-ALLand AL-PUBdatasets, a collection of images of handwritten ancient Greek characters that were labeled by volunteers through the Ancient Lives online citizen science project. We find our best-performing machine learning model to be 92.57% accurate compared to Tesseract OCR’s 11.15%. Following our analysis, we present a brief examination of our models’ shortcomings, introduce the publicly-available AL-PUBdataset, and, describe Theia, a web-based tool that democratizes our machine learning models for public use. We conclude by discussing the promise of our findings for advancing research at the intersection of machine learning, manuscript transcription, and the digital humanities.
DOI: --
发表时间: 2002-09
期刊: --
影响因子: --
作者:
Xiaojin Zhu;Zoubin Ghahramani
通讯作者: Xiaojin Zhu;Zoubin Ghahramani
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者:
R. Grayson
通讯作者: R. Grayson
DOI: --
发表时间: 2010
期刊: Electronic imaging
影响因子: --
作者:
Markus Diem;Robert Sablatnig
通讯作者: Robert Sablatnig
影响因子: 0.9
作者:
Nurshazlyn M. Aszemi;P. Dominic
通讯作者: P. Dominic
数字纸莎草学 I:方法、工具和趋势
DOI: 10.1515/9783110547474
发表时间: 2017
期刊: PLoS ONE
影响因子: 3.7
作者:
N. Reggiani
通讯作者: N. Reggiani