Labeling Poststorm Coastal Imagery for Machine Learning: Measurement of Interrater Agreement

Labeling Poststorm Coastal Imagery for Machine Learning: Measurement of Interrater Agreement
复制标题

DOI:
10.1029/2021ea001896
复制
发表时间:
2021-09-01
影响因子:
3.1
通讯作者:
Williams, Hannah E.
Williams, Hannah E.
中科院分区:
地球科学3区
文献类型:
--
作者:
Goldstein, Evan B.;Buscombe, Daniel;Williams, Hannah E.

文献摘要

被引文献

相似文献

使用监督机器学习(ML)对图像进行分类依赖于标记的训练数据类或文本描述,例如,与每个图像相关联。数据驱动的模型与用于训练的数据一样好,这指出了开发具有预测技能的ML模型的高质量标记数据的重要性。标记数据通常是一个耗时的手动过程。在这里,我们研究了标记数据的过程,特别关注飓风影响美国大西洋和墨西哥湾沿岸后捕获的沿海航空图像。图像数据集是风暴影响和海岸变化的丰富观测记录,但图像需要标记以使这些信息易于获取。我们创建了一个在线界面,为标签者提供一系列图像和一组固定的问题。总共有1600张图片被至少两名或多达七名海岸科学家标记。我们使用结果数据集来调查标注者一致性:标注者标记每个图像相似的程度。当向标注者提出的问题相对简单时,当标注者提供用户手册时,当图像较小时,用百分比一致性和Krippendorff的alpha来评估的判读者协议分数更高。解释器协议的实验表明,多个标注器对于理解机器学习研究中标注数据的不确定性有好处。飓风和风暴过后,从飞机上拍摄的照片可以用来观察海岸是如何受到影响的。一次飞行可能会拍摄数千张照片。如果计算机可以自动分析这些图片,那么人们就不需要逐一查看了。为了教计算机分析图像,我们需要许多图片和许多标签来描述每张图片中可见的内容。但我们从哪里得到这些标签呢?通常情况下,沿海科学家通过将图片分类到文件夹中或将代码输入电子表格来标记图片。但是每个海岸科学家都以同样的方式给图片贴上标签吗?一些标签问题很容易回答,科学家们也大多同意(“这张图全是水吗?”)。其他标签问题更难回答,也会引起分歧(“建筑物受损了吗?”)。这篇论文是关于科学家在标记相同图片时的一致程度,以及我们如何提高科学家之间的一致。我们尝试了一些实验,并就如何提高一致性提出了一些想法。我们建议写非常清晰的问题,使用较小的图像,并有一个全面的手册。事实证明,拥有一本带有示例的手册——并阅读手册!真的有帮助。
Classifying images using supervised machine learning (ML) relies on labeled training data-classes or text descriptions, for example, associated with each image. Data-driven models are only as good as the data used for training, and this points to the importance of high-quality labeled data for developing a ML model that has predictive skill. Labeling data is typically a time-consuming, manual process. Here, we investigate the process of labeling data, with a specific focus on coastal aerial imagery captured in the wake of hurricanes that affected the Atlantic and Gulf Coasts of the United States. The imagery data set is a rich observational record of storm impacts and coastal change, but the imagery requires labeling to render that information accessible. We created an online interface that served labelers a stream of images and a fixed set of questions. A total of 1,600 images were labeled by at least two or as many as seven coastal scientists. We used the resulting data set to investigate interrater agreement: the extent to which labelers labeled each image similarly. Interrater agreement scores, assessed with percent agreement and Krippendorff's alpha, are higher when the questions posed to labelers are relatively simple, when the labelers are provided with a user manual, and when images are smaller. Experiments in interrater agreement point toward the benefit of multiple labelers for understanding the uncertainty in labeling data for machine learning research.Plain Language Summary After hurricanes and storms, pictures taken from a plane can be used to observe how the coast was impacted. A single flight might take thousands of pictures. If a computer could automatically analyze the pictures, then a person would not need to look at them one-byone. To teach a computer to analyze images, we need many pictures and many labels that describe what is visible in each picture. But where do we get those labels? Typically, a coastal scientist labels the pictures by sorting them into folders or typing codes into a spreadsheet. But does every coastal scientist label pictures the same way? Some labeling questions are easy to answer, and scientists mostly agree ("Is this image all water?"). Other labeling questions are harder to answer, and cause disagreement ("Was there damage to buildings?"). This paper is about how well scientists agree when labeling the same pictures, and how we can improve agreement among scientists. We try some experiments and offer a few ideas on how to improve agreement. We suggest writing very clear questions, using smaller images, and having a comprehensive manual. It turns out that having a manual with examples-and reading the manual!really helps.