The National Eutrophication Survey: lake characteristics and historical nutrient concentrations

The National Eutrophication Survey: lake characteristics and historical nutrient concentrations
复制标题

DOI:
10.5194/essd-10-81-2018
复制
发表时间:
2017-07
影响因子:
11.4
通讯作者:
J. Stachelek;C. Ford;D. Kincaid;Katelyn B. S. King;H. Miller;R. Nagelkirk
J. Stachelek;C. Ford;D. Kincaid;Katelyn B. S. King;H. Miller;R. Nagelkirk
中科院分区:
地球科学1区
文献类型:
--
作者:
J. Stachelek;C. Ford;D. Kincaid;Katelyn B. S. King;H. Miller;R. Nagelkirk

文献摘要

被引文献

相似文献

抽象。历史生态调查作为一个基线,并为当代研究提供背景,但许多这些记录没有保存的方式,以确保其长期可用性。国家富营养化调查(内斯)数据库目前仅以原始报告(PDF文件)的扫描形式提供,没有嵌入字符信息。这限制了它的可搜索性,机器可读性以及当前和未来科学家系统评估其内容的能力。内斯的数据是由美国环境保护署在1972年至1975年期间收集的,作为调查淡水湖泊和水库富营养化的努力的一部分。虽然有几项研究已经手动转录了数据库的一小部分,以支持特定的研究,但还没有系统地尝试转录和保存整个数据库。在这里,我们使用自动光学字符识别和手动质量保证程序的组合,使这些数据可用于分析。发现光学字符识别协议的性能与原始文件质量(清晰度)的变化有关。对于四份存档扫描报告中的每一份,我们的质量保证协议发现错误率在5.9%到17%之间。我们的方法的目标是通过将手工输入数据与数字转录技术相结合,在效率和数据质量之间取得平衡。完成的数据库包含美国周边约800个湖泊的物理特征、水文和水质信息(Stachelek et al.,2017,https://doi.org/10.5063/F1639MVD)。最终,该数据库可以与最近的研究相结合,以生成对美国大陆水质趋势和空间变化的荟萃分析。
Abstract. Historical ecological surveys serve as a baseline and provide context for contemporary research, yet many of these records are not preserved in a way that ensures their long-term usability. The National Eutrophication Survey (NES) database is currently only available as scans of the original reports (PDF files) with no embedded character information. This limits its searchability, machine readability, and the ability of current and future scientists to systematically evaluate its contents. The NES data were collected by the US Environmental Protection Agency between 1972 and 1975 as part of an effort to investigate eutrophication in freshwater lakes and reservoirs. Although several studies have manually transcribed small portions of the database in support of specific studies, there have been no systematic attempts to transcribe and preserve the database in its entirety. Here we use a combination of automated optical character recognition and manual quality assurance procedures to make these data available for analysis. The performance of the optical character recognition protocol was found to be linked to variation in the quality (clarity) of the original documents. For each of the four archival scanned reports, our quality assurance protocol found an error rate between 5.9 and 17 % . The goal of our approach was to strike a balance between efficiency and data quality by combining entry of data by hand with digital transcription technologies. The finished database contains information on the physical characteristics, hydrology, and water quality of about 800 lakes in the contiguous US ( Stachelek et al. , 2017 , https://doi.org/10.5063/F1639MVD ). Ultimately, this database could be combined with more recent studies to generate meta-analyses of water quality trends and spatial variation across the continental US.