Unlocking digital archives: cross-disciplinary perspectives on AI and born-digital data.

Unlocking digital archives: cross-disciplinary perspectives on AI and born-digital data.
复制标题

DOI:
10.1007/s00146-021-01367-x
复制
发表时间:
2022
期刊:
影响因子:
3
通讯作者:
Caputo A
Caputo A
中科院分区:
其他
文献类型:
--
作者:
Jaillant L;Caputo A

文献摘要

参考文献

相似文献

本文由计算机科学家和数字人文主义者共同撰写,探讨了文化遗产机构在数字时代面临的挑战,这些挑战导致了绝大多数非数字档案馆的关闭。它特别关注历史学家、文学学者和其他人文学者使用的图书馆、博物馆和档案馆等文化组织。由于隐私、版权、商业和技术问题,文化组织持有的大多数数字记录都无法访问。即使是公开的数字化数据(如网络档案),用户也往往需要亲自前往大英图书馆或法国国家图书馆等资料库查阅网页。有了足够的样本数据来学习和训练他们的模型,人工智能,更具体地说是机器学习算法,通过学习执行复杂的人工任务,提供了改善和简化数字档案访问的机会。从为档案搜索提供智能支持到自动化繁琐耗时的任务,这些都有所不同。在本文中,我们将重点关注敏感性审查作为一种实用的解决方案,以解锁数字档案,使档案机构能够提供非敏感信息。这种使档案更容易获得的承诺并不是没有潜在陷阱和风险的警告:固有的错误,使算法难以理解的“黑箱”方法,以及与偏见,虚假或部分信息相关的风险。我们的中心论点是,人工智能可以实现其承诺,使数字档案收藏更容易获得,但它也带来了新的挑战-特别是在伦理方面。最后,我们坚持认为,在使数字档案更容易获取的过程中,公平、问责和透明至关重要。
Co-authored by a Computer Scientist and a Digital Humanist, this article examines the challenges faced by cultural heritage institutions in the digital age, which have led to the closure of the vast majority of born-digital archival collections. It focuses particularly on cultural organizations such as libraries, museums and archives, used by historians, literary scholars and other Humanities scholars. Most born-digital records held by cultural organizations are inaccessible due to privacy, copyright, commercial and technical issues. Even when born-digital data are publicly available (as in the case of web archives), users often need to physically travel to repositories such as the British Library or the Bibliothèque Nationale de France to consult web pages. Provided with enough sample data from which to learn and train their models, AI, and more specifically machine learning algorithms, offer the opportunity to improve and ease the access to digital archives by learning to perform complex human tasks. These vary from providing intelligent support for searching the archives to automate tedious and time-consuming tasks.  In this article, we focus on sensitivity review as a practical solution to unlock digital archives that would allow archival institutions to make non-sensitive information available. This promise to make archives more accessible does not come free of warnings for potential pitfalls and risks: inherent errors, "black box" approaches that make the algorithm inscrutable, and risks related to bias, fake, or partial information. Our central argument is that AI can deliver its promise to make digital archival collections more accessible, but it also creates new challenges - particularly in terms of ethics. In the conclusion, we insist on the importance of fairness, accountability and transparency in the process of making digital archives more accessible.
DOI: 10.1080/01576895.2018.1502088
发表时间: 2019-01-01
影响因子: 0.3
作者:
Rolan, Gregory;Humphries, Glen;Stuart, Katharine
通讯作者: Stuart, Katharine
DOI: 10.1007/s11023-020-09517-8
发表时间: 2020-02-01
期刊: MINDS AND MACHINES
影响因子: 7.4
作者:
Hagendorff, Thilo
通讯作者: Hagendorff, Thilo
DOI: 10.1038/s42256-019-0088-2
发表时间: 2019-09-01
影响因子: 23.8
作者:
Jobin, Anna;Ienca, Marcello;Vayena, Effy
通讯作者: Vayena, Effy
DOI: 10.1145/1749603.1749605
发表时间: 2010-06-01
影响因子: 16.6
作者:
Fung, Benjamin C. M.;Wang, Ke;Yu, Philip S.
通讯作者: Yu, Philip S.
DOI: 10.7202/303466ar
发表时间: 1975-01-01
期刊: REVISTA DE HISTORIA DE AMERICA
影响因子: --
作者:
DUMONTJOHNSON, M
通讯作者: DUMONTJOHNSON, M