Fast, scalable, and automated identification of articles for biodiversity and macroecological datasets

Fast, scalable, and automated identification of articles for biodiversity and macroecological datasets
复制标题

DOI:
10.1111/geb.13219
复制
发表时间:
2020-11-19
影响因子:
6.4
通讯作者:
Freeman, Robin
Freeman, Robin
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Cornford, Richard;Deinet, Stefanie;Freeman, Robin

文献摘要

被引文献

相似文献

目的了解大尺度的生态模式和过程是必要的,如果我们要减轻的后果,生物多样性退化。然而,这种分析需要大量的数据集,目前的数据整理方法可能很慢,涉及大量的人工输入。鉴于科学出版物的快速和不断增长的速度,在数十万篇文章中手动识别数据源是一项重大挑战,这可能会在生态数据库的生成中造成瓶颈。创新在这里,我们展示了使用一般的文本分类方法来识别相关的生物多样性文章。我们将其应用到两个免费提供的示例数据库,生命星球数据库和数据库的预测(预测响应变化的陆地系统中的生态多样性)项目,这两个都是重要的生物多样性指标的基础。我们评估了基于逻辑回归(LR)和卷积神经网络的机器学习分类器,并确定了影响分类性能的文本处理工作流程的各个方面。主要结论我们最好的分类器可以区分相关和非相关文章,准确率超过90%。使用现成的摘要和标题或单独使用摘要比单独使用标题产生更好的结果。LR和神经网络模型表现相似。至关重要的是,我们表明,与当前的手动协议相比,在现实世界的搜索结果上部署此类模型可以显着提高潜在相关论文的恢复率。此外,我们的研究结果表明,给定100篇相关论文的适度初始样本,可以通过基于目标文献搜索迭代更新训练文本来快速生成高性能分类器。这些研究结果清楚地表明,文本挖掘方法用于构建和增强生态数据集的有用性,这些技术的更广泛应用有可能使更广泛的大规模分析受益。我们提供了源代码和示例,可用于为其他数据集创建新的分类器。
Aim Understanding broad-scale ecological patterns and processes is necessary if we are to mitigate the consequences of anthropogenically driven biodiversity degradation. However, such analyses require large datasets and current data collation methods can be slow, involving extensive human input. Given rapid and ever-increasing rates of scientific publication, manually identifying data sources among hundreds of thousands of articles is a significant challenge, which can create a bottleneck in the generation of ecological databases.Innovation Here, we demonstrate the use of general, text-classification approaches to identify relevant biodiversity articles. We apply this to two freely available example databases, the Living Planet Database and the database of the PREDICTS (Projecting Responses of Ecological Diversity in Changing Terrestrial Systems) project, both of which underpin important biodiversity indicators. We assess machine-learning classifiers based on logistic regression (LR) and convolutional neural networks, and identify aspects of the text-processing workflow that influence classification performance.Main conclusions Our best classifiers can distinguish relevant from non-relevant articles with over 90% accuracy. Using readily available abstracts and titles or abstracts alone produces significantly better results than using titles alone. LR and neural network models performed similarly. Crucially, we show that deploying such models on real-world search results can significantly increase the rate at which potentially relevant papers are recovered compared to a current manual protocol. Furthermore, our results indicate that, given a modest initial sample of 100 relevant papers, high-performing classifiers could be generated quickly through iteratively updating the training texts based on targeted literature searches. These findings clearly demonstrate the usefulness of text-mining methods for constructing and enhancing ecological datasets, and wider application of these techniques has the potential to benefit large-scale analyses more broadly. We provide source code and examples that can be used to create new classifiers for other datasets.