Computational metadata generation methods for biological specimen image collections

Computational metadata generation methods for biological specimen image collections
复制标题

DOI:
10.1007/s00799-022-00342-1
复制
发表时间:
2022-11-23
影响因子:
1.5
通讯作者:
Greenberg, Jane
Greenberg, Jane
中科院分区:
其他
文献类型:
--
作者:
Karnani, Kevin;Pepper, Joel;Greenberg, Jane

文献摘要

被引文献

相似文献

元数据是研究人员寻求将机器学习(ML)应用于可在线找到的大量数字化生物标本的关键数据源。不幸的是,相关的元数据往往是稀疏的,有时是错误的。本文扩展了以前的研究与伊利诺伊州自然历史调查(INHS)收集(7244标本图像),使用计算方法来分析图像质量,然后自动生成22元数据属性代表的图像质量和形态特征的标本。在这里报告的研究中,我们展示了我们的初步工作的扩展到威斯康星州动物博物馆(UWZM)收集(4155标本图像)。此外,我们以四种方式增强了我们的计算方法:(1)增强训练集,(2)应用对比度增强,(3)放大小对象,以及(4)改进我们的处理逻辑。这些新方法将我们的总体错误率从4.6%提高到1.1%。这些增强还允许我们计算额外的17个基于图像的元数据属性。新的元数据属性提供了补充功能和信息,也可用于分析和分类的鱼类标本。这些新功能的示例包括凸面积、偏心率、周长、偏斜等。新改进的过程在时间和劳动力成本以及准确性方面进一步优于人类,为利用ML数字化标本提供了一种新的解决方案。这项研究证明了计算方法的能力,以提高与数以万计的数字化标本存储在世界各地的开放存取存储库,通过生成准确的和有价值的元数据,这些存储库相关的数字图书馆服务。
Metadata is a key data source for researchers seeking to apply machine learning (ML) to the vast collections of digitized biological specimens that can be found online. Unfortunately, the associated metadata is often sparse and, at times, erroneous. This paper extends previous research conducted with the Illinois Natural History Survey (INHS) collection (7244 specimen images) that uses computational approaches to analyze image quality, and then automatically generates 22 metadata properties representing the image quality and morphological features of the specimens. In the research reported here, we demonstrate the extension of our initial work to University of the Wisconsin Zoological Museum (UWZM) collection (4155 specimen images). Further, we enhance our computational methods in four ways: (1) augmenting the training set, (2) applying contrast enhancement, (3) upscaling small objects, and (4) refining our processing logic. Together these new methods improved our overall error rates from 4.6 to 1.1%. These enhancements also allowed us to compute an additional set of 17 image-based metadata properties. The new metadata properties provide supplemental features and information that may also be used to analyze and classify the fish specimens. Examples of these new features include convex area, eccentricity, perimeter, skew, etc. The newly refined process further outperforms humans in terms of time and labor cost, as well as accuracy, providing a novel solution for leveraging digitized specimens with ML. This research demonstrates the ability of computational methods to enhance the digital library services associated with the tens of thousands of digitized specimens stored in open-access repositories world-wide by generating accurate and valuable metadata for those repositories.