Tweet Classification Toward Twitter-Based Disease Surveillance: New Data, Methods, and Evaluations

Tweet Classification Toward Twitter-Based Disease Surveillance: New Data, Methods, and Evaluations
复制标题

DOI:
10.2196/12783
复制
发表时间:
2019-02-20
影响因子:
7.4
通讯作者:
Aramaki, Eiji
Aramaki, Eiji
中科院分区:
医学2区
文献类型:
--
作者:
Wakamiya, Shoko;Morita, Mizuki;Aramaki, Eiji

文献摘要

被引文献

相似文献

背景:Web上的医学和临床相关信息量正在增加。在现有的不同类型的信息中,直接从人们那里获得的基于社交媒体的数据特别有价值,并正在引起人们的极大关注。为了鼓励利用社交媒体数据进行医学自然语言处理(NLP)研究,第13届NII测试床和信息访问研究社区(NTCIR-13)Web文档的医学自然语言处理(MedWeb)在跨语言和多标签语料库中提供伪Twitter消息,覆盖3种语言(日语、英语和中文),并附有8个症状标签(如感冒、发烧和流感)。然后,参与者将每个tweet分为两类:包含患者症状的tweet和不包含患者症状的tweet。目的:本研究旨在通过讨论展示参与日语子任务、英语子任务和汉语子任务的小组的结果,以澄清医学NLP领域需要解决的问题。综上所述,8组(19个系统)参与了日语子任务,4组(12个系统)参与了英语子任务,2组(6个系统)参与了汉语子任务。每个子任务总共构建了2个基线系统。参与者和基线系统的性能进行了评估,使用精确的匹配准确率,F-测量的基础上的精度和召回率,和Hamming loss.Results:最好的系统达到了精确的0.880匹配准确率,0.920 F-测量,和0.019汉明损失。日语子任务的匹配准确度、F测量和汉明损失的平均值分别为0.720、0.820和0.051;英语子任务的平均值分别为0.770、0.850和0.037;汉语子任务的平均值分别为0.810、0.880和0.032。本文介绍并讨论了参与NTCIR-13 MedWeb任务的系统的性能。由于MedWeb任务设置可以形式化为文本的事实化,因此该任务的实现可以直接应用于实际的临床应用。
Background: The amount of medical and clinical-related information on the Web is increasing. Among the different types of information available, social media-based data obtained directly from people are particularly valuable and are attracting significant attention. To encourage medical natural language processing (NLP) research exploiting social media data, the 13th NII Testbeds and Community for Information access Research (NTCIR-13) Medical natural language processing for Web document (MedWeb) provides pseudo-Twitter messages in a cross-language and multi-label corpus, covering 3 languages (Japanese, English, and Chinese) and annotated with 8 symptom labels (such as cold, fever, and flu). Then, participants classify each tweet into 1 of the 2 categories: those containing a patient's symptom and those that do not.Objective: This study aimed to present the results of groups participating in a Japanese subtask, English subtask, and Chinese subtask along with discussions, to clarify the issues that need to be resolved in the field of medical NLP.Methods: In summary, 8 groups (19 systems) participated in the Japanese subtask, 4 groups (12 systems) participated in the English subtask, and 2 groups (6 systems) participated in the Chinese subtask. In total, 2 baseline systems were constructed for each subtask. The performance of the participant and baseline systems was assessed using the exact match accuracy, F-measure based on precision and recall, and Hamming loss.Results: The best system achieved exactly 0.880 match accuracy, 0.920 F-measure, and 0.019 Hamming loss. The averages of match accuracy, F-measure, and Hamming loss for the Japanese subtask were 0.720, 0.820, and 0.051; those for the English subtask were 0.770, 0.850, and 0.037; and those for the Chinese subtask were 0.810, 0.880, and 0.032, respectively.Conclusions: This paper presented and discussed the performance of systems participating in the NTCIR-13 MedWeb task. As the MedWeb task settings can be formalized as the factualization of text, the achievement of this task could be directly applied to practical clinical applications.