No More 404s: Predicting Referenced Link Rot in Scholarly Articles for Pro-Active Archiving
No More 404s: Predicting Referenced Link Rot in Scholarly Articles for Pro-Active Archiving
复制标题
不再出现 404:预测学术文章中的引用链接失效以进行主动归档
DOI:
10.1145/2756406.2756940
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
R. Tobin
中科院分区:
文献类型:
--
作者:
Ke Zhou;Claire Grover;Martin Klein;R. Tobin
The citation of resources is a fundamental part of scholarly discourse. Due to the popularity of the web, there is an increasing trend for scholarly articles to reference web resources (e.g. software, data). However, due to the dynamic nature of the web, the referenced links may become inaccessible ('rotten') sometime after publication, returning a "404 Not Found" HTTP error. In this paper we first present some preliminary findings of a study of the persistence and availability of web resources referenced from papers in a large-scale scholarly repository. We reaffirm previous research that link rot is a serious problem in the scholarly world and that current web archives do not always preserve all rotten links. Therefore, a more pro-active archival solution needs to be developed to further preserve web content referenced in scholarly articles. To this end, we propose to apply machine learning techniques to train a link rot predictor for use by an archival framework to prioritise pro-active archiving of links that are more likely to be rotten. We demonstrate that we can obtain a fairly high link rot prediction AUC (0.72) with only a small set of features. By simulation, we also show that our prediction framework is more effective than current web archives for preserving links that are likely to be rotten. This work has a potential impact for the scholarly world where publishers can utilise this framework to prioritise the archiving of links for digital preservation, especially when there is a large quantity of links to be archived.