Computer vision and machine learning approaches for metadata enrichment to improve searchability of historical newspaper collections

Computer vision and machine learning approaches for metadata enrichment to improve searchability of historical newspaper collections
复制标题

DOI:
10.1108/jd-01-2022-0029
复制
发表时间:
2023-02
影响因子:
2.1
通讯作者:
Dilawar Ali;Kenzo Milleville;S. Verstockt;N. van de Weghe;Sally Chambers;Julie M. Birkholz
Dilawar Ali;Kenzo Milleville;S. Verstockt;N. van de Weghe;Sally Chambers;Julie M. Birkholz
中科院分区:
管理学3区
文献类型:
--
作者:
Dilawar Ali;Kenzo Milleville;S. Verstockt;N. van de Weghe;Sally Chambers;Julie M. Birkholz

文献摘要

相似文献

目的历史报纸收藏提供了丰富的关于过去的信息。虽然这些馆藏的数字化大大提高了它们的可访问性,但大部分数字化的历史报纸馆藏,如比利时皇家图书馆的历史报纸馆藏,还不能在文章级别进行搜索。然而,最近基于人工智能的研究方法的发展,如文档布局分析,有可能进一步丰富元数据,以提高这些历史报纸收藏的可搜索性。本文旨在对上述问题进行探讨。在本文中,作者探讨了如何使用现有的计算机视觉和机器学习方法来改善对数字化历史报纸的访问。为此,作者提出了一个工作流程,使用计算机视觉和机器学习方法(1)使用文档布局分析提供对数字化历史报纸馆藏的文章级别访问,(2)提取特定类型的文章(例如1938年《Le people》的文学补充),(3)使用(非)监督分类方法进行图像相似性分析,(4)执行命名实体识别(NER)将提取的信息链接到开放数据。结果表明,所提出的工作流程提高了数字化历史报纸的可及性和可检索性,并有助于数字人文学科研究的语料库建设。基于人工智能的方法可以自动提取特征,对相似图像进行聚类,并动态链接相关文章。原创性/价值提议的工作流程能够自动提取文章,包括检测特定类型的文章,如专栏或文学补充。这对人文研究人员来说尤其有价值,因为它提高了这些集合的可搜索性,并使语料库能够围绕特定主题构建。通过在线工具(https://tw06v072.ugent.be/kbr/)展示了对KBR数字化报纸的文章级访问和改进的可搜索性。
PurposeHistorical newspaper collections provide a wealth of information about the past. Although the digitization of these collections significantly improves their accessibility, a large portion of digitized historical newspaper collections, such as those of KBR, the Royal Library of Belgium, are not yet searchable at article-level. However, recent developments in AI-based research methods, such as document layout analysis, have the potential for further enriching the metadata to improve the searchability of these historical newspaper collections. This paper aims to discuss the aforementioned issue.Design/methodology/approachIn this paper, the authors explore how existing computer vision and machine learning approaches can be used to improve access to digitized historical newspapers. To do this, the authors propose a workflow, using computer vision and machine learning approaches to (1) provide article-level access to digitized historical newspaper collections using document layout analysis, (2) extract specific types of articles (e.g. feuilletons – literary supplements from Le Peuple from 1938), (3) conduct image similarity analysis using (un)supervised classification methods and (4) perform named entity recognition (NER) to link the extracted information to open data.FindingsThe results show that the proposed workflow improves the accessibility and searchability of digitized historical newspapers, and also contributes to the building of corpora for digital humanities research. The AI-based methods enable automatic extraction of feuilletons, clustering of similar images and dynamic linking of related articles.Originality/valueThe proposed workflow enables automatic extraction of articles, including detection of a specific type of article, such as a feuilleton or literary supplement. This is particularly valuable for humanities researchers as it improves the searchability of these collections and enables corpora to be built around specific themes. Article-level access to, and improved searchability of, KBR's digitized newspapers are demonstrated through the online tool (https://tw06v072.ugent.be/kbr/).