A Large Dataset of Historical Japanese Documents with Complex Layouts

A Large Dataset of Historical Japanese Documents with Complex Layouts
复制标题

DOI:
10.1109/cvprw50498.2020.00282
复制
发表时间:
2020-04
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Zejiang Shen;Kaixuan Zhang;Melissa Dell
Zejiang Shen;Kaixuan Zhang;Melissa Dell
中科院分区:
其他
文献类型:
--
作者:
Zejiang Shen;Kaixuan Zhang;Melissa Dell

文献摘要

被引文献

相似文献

基于深度学习的自动文档布局分析和内容提取方法有可能大规模地解锁困在历史文档中的丰富信息。一个主要障碍是缺乏用于训练稳健模型的大型数据集。特别是,亚洲语言的训练数据很少。为此,我们提出了一个复杂布局的大型日本历史文献数据集HJDataset。它包含超过25万个7种类型的布局元素注释。除了内容区域的边界框和掩码之外,它还包括布局元素的层次结构和读取顺序。该数据集是由人工和机器共同构建的。提出了一种基于半规则的布局元素提取方法,并由人工检验员对结果进行检验。由此产生的大规模数据集用于使用最先进的深度学习模型为文本区域检测提供基线性能分析。我们展示了数据集在现实世界文档数字化任务中的有用性。该数据集可在https://dell-research-harvard.github.io/HJDataset/上获得。
Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for training robust models. In particular, little training data exist for Asian languages. To this end, we present HJDataset, a Large Dataset of Historical Japanese Documents with Complex Layouts. It contains over 250,000 layout element annotations of seven types. In addition to bounding boxes and masks of the content regions, it also includes the hierarchical structures and reading orders for layout elements. The dataset is constructed using a combination of human and machine efforts. A semi-rule based method is developed to extract the layout elements, and the results are checked by human inspectors. The resulting large-scale dataset is used to provide baseline performance analyses for text region detection using state-of-the-art deep learning models. And we demonstrate the usefulness of the dataset on real-world document digitization tasks. The dataset is available at https://dell-research-harvard.github.io/HJDataset/.