RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media

RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media
复制标题

DOI:
10.48550/arxiv.2210.06331
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Somin Wadhwa;Vivek Khetan;Silvio Amir;Byron Wallace
Somin Wadhwa;Vivek Khetan;Silvio Amir;Byron Wallace
中科院分区:
其他
文献类型:
--
作者:
Somin Wadhwa;Vivek Khetan;Silvio Amir;Byron Wallace

文献摘要

被引文献

相似文献

我们介绍了Reddit健康在线谈话(RedHOT),这是一个语料库,包含来自Reddit的22,000个注释丰富的社交媒体帖子,涵盖24种健康状况。注释包括与医疗索赔、个人经历和问题相对应的跨度划分。我们收集已识别索赔的其他粒度注释。具体而言,我们标记描述患者人群、干预措施和结局(PIO元素)的片段。使用这个语料库,我们介绍了检索与社交媒体上的给定声明相关的可信证据的任务。我们提出了一种新的方法来自动获得(噪声)监督这项任务,我们用它来训练一个密集的检索模型,这优于基线模型。由医生执行的检索结果的手动评估表明,虽然我们的系统性能是有前途的,有相当大的改进空间。我们发布了收集的所有注释(和脚本组装数据集),以及在本文中重现结果所需的所有代码:https://sominw.com/redhot。
We present Reddit Health Online Talk (RedHOT), a corpus of 22,000 richly annotated social media posts from Reddit spanning 24 health conditions. Annotations include demarcations of spans corresponding to medical claims, personal experiences, and questions.We collect additional granular annotations on identified claims.Specifically, we mark snippets that describe patient Populations, Interventions, and Outcomes (PIO elements) within these. Using this corpus, we introduce the task of retrieving trustworthy evidence relevant to a given claim made on social media. We propose a new method to automatically derive (noisy) supervision for this task which we use to train a dense retrieval model; this outperforms baseline models. Manual evaluation of retrieval results performed by medical doctors indicate that while our system performance is promising, there is considerable room for improvement.We release all annotations collected (and scripts to assemble the dataset), and all code necessary to reproduce the results in this paper at: https://sominw.com/redhot.