Garbage in, Garbage Out: Data Collection, Quality Assessment and Reporting Standards for Social Media Data Use in Health Research, Infodemiology and Digital Disease Detection.

Garbage in, Garbage Out: Data Collection, Quality Assessment and Reporting Standards for Social Media Data Use in Health Research, Infodemiology and Digital Disease Detection.
复制标题

DOI:
10.2196/jmir.4738
复制
发表时间:
2016-02-26
影响因子:
7.4
通讯作者:
Emery S
Emery S
中科院分区:
医学2区
文献类型:
--
作者:
Kim Y;Huang J;Emery S

文献摘要

被引文献

相似文献

社交媒体已经改变了通信的格局。人们越来越多地通过网络和社交媒体获取新闻和健康信息。社交媒体平台也为卫生研究(包括信息流行病学、信息监测和数字疾病检测)提供了丰富的观测数据的新来源。虽然使用社交数据的研究数量正在迅速增长,但这些研究中很少有透明地概述他们收集、过滤和报告这些数据的方法。应用于社交数据的关键词和搜索过滤器形成了研究人员可以观察人们就给定话题进行交流的内容和方式的透镜。如果没有正确聚焦的镜头,研究结论可能会有偏见或误导。报告数据源和质量的标准是必要的,这样数据科学家和社交媒体研究的消费者就可以评估和比较各种研究的方法和结果。我们的目标是开发和应用一个社交媒体数据收集和质量评估的框架,并提出一个报告标准,研究人员和审稿人可以使用它来评估和比较研究中社交数据的质量。我们提出了一个概念性框架,包括收集社交媒体数据的三个主要步骤:开发、应用和验证搜索过滤器。这个框架基于两个标准:检索精度(检索的数据中有多少是相关的)和检索召回率(检索的相关数据中有多少)。然后,我们讨论了检索精度和召回率的估计依赖于准确的人类编码和完整的数据收集的两个条件,以及如何在偏离这两个理想条件的情况下计算这些统计数据。然后,我们将该框架应用于一个真实世界的示例,该示例使用从Twitter消防水带收集的大约400万条与烟草相关的tweet。我们开发并应用了一个搜索过滤器,根据三个关键字类别:设备、品牌和行为,从存档中检索与电子烟相关的推文。搜索过滤器从存档中检索了82205条与电子烟相关的推文,并进行了验证。所有病例的检索精度均在95%以上。假设理想条件(没有人为编码错误和完整的数据收集),检索召回率为86%,未检索的消息无法存档时为75%,86%假设编码员没有假阴性错误,93%允许人为编码员的假阴性和假阳性错误。本文提出了一个社会数据过滤和质量评估的概念框架,解决了几个共同的挑战,并朝着建立报告社会数据的标准迈进。研究人员应该清楚地描述数据的来源,如何访问和收集数据,以及搜索过滤器的构建过程以及如何计算检索精度和召回率。建议的框架可以适用于其他公共社交媒体平台。
Social media have transformed the communications landscape. People increasingly obtain news and health information online and via social media. Social media platforms also serve as novel sources of rich observational data for health research (including infodemiology, infoveillance, and digital disease detection detection). While the number of studies using social data is growing rapidly, very few of these studies transparently outline their methods for collecting, filtering, and reporting those data. Keywords and search filters applied to social data form the lens through which researchers may observe what and how people communicate about a given topic. Without a properly focused lens, research conclusions may be biased or misleading. Standards of reporting data sources and quality are needed so that data scientists and consumers of social media research can evaluate and compare methods and findings across studies. We aimed to develop and apply a framework of social media data collection and quality assessment and to propose a reporting standard, which researchers and reviewers may use to evaluate and compare the quality of social data across studies. We propose a conceptual framework consisting of three major steps in collecting social media data: develop, apply, and validate search filters. This framework is based on two criteria: retrieval precision (how much of retrieved data is relevant) and retrieval recall (how much of the relevant data is retrieved). We then discuss two conditions that estimation of retrieval precision and recall rely on—accurate human coding and full data collection—and how to calculate these statistics in cases that deviate from the two ideal conditions. We then apply the framework on a real-world example using approximately 4 million tobacco-related tweets collected from the Twitter firehose. We developed and applied a search filter to retrieve e-cigarette–related tweets from the archive based on three keyword categories: devices, brands, and behavior. The search filter retrieved 82,205 e-cigarette–related tweets from the archive and was validated. Retrieval precision was calculated above 95% in all cases. Retrieval recall was 86% assuming ideal conditions (no human coding errors and full data collection), 75% when unretrieved messages could not be archived, 86% assuming no false negative errors by coders, and 93% allowing both false negative and false positive errors by human coders. This paper sets forth a conceptual framework for the filtering and quality evaluation of social data that addresses several common challenges and moves toward establishing a standard of reporting social data. Researchers should clearly delineate data sources, how data were accessed and collected, and the search filter building process and how retrieval precision and recall were calculated. The proposed framework can be adapted to other public social media platforms.