A benchmark dataset and case study for Chinese medical question intent classification

A benchmark dataset and case study for Chinese medical question intent classification
复制标题

中文医学问题意图分类基准数据集及案例研究

DOI:
10.1186/s12911-020-1122-3
复制
发表时间:
2020-07-09
影响因子:
3.5
通讯作者:
Wei, Ming
Wei, Ming
中科院分区:
医学3区
文献类型:
--
作者:
Chen, Nan;Su, Xiangdong;Wei, Ming

文献摘要

被引文献

相似文献

背景医学问答系统必须准确理解用户提问的意图,才能提供满意的答案。对于医疗意图分类,它需要高质量的数据集来以监督的方式训练深度学习方法。目前,中国医疗意图分类没有公开的数据集,其他领域的数据集也不适用于医疗QA系统。为了解决这个问题,我们使用来自医疗问答网站的问题构建了一个中国医疗意图数据集(CMID)。在此基础上,我们比较了四个意图分类模型CMID使用caseStudy.MethodsThe CMID中的问题是从几个医疗QA网站。意图标注标准是由医学专家制定的,包括用户意图的四种类型和36个子类型。除了Intent标签,CMID还提供了两种类型的附加信息,包括分词和命名实体。我们使用众包的方式来注释每个中国医学问题的意图信息。使用Jieba和经过良好训练的Lattice-LSTM模型获得分词和命名实体。我们加载了一个由53万个词组成的中文医学词典进行分词,以获得更准确的结果。我们还选择了四个流行的基于深度学习的模型,并比较其性能的意图分类CMID.ResultsThe最终CMID包含12,000中国医疗问题,并组织在JSON格式。每个问题都标记了意图、分词和命名实体信息。对问题长度、实体个数、以及问题类型等信息也进行了详细的分析.在Fast Text、TextCNN、TextRNN和TextGCN中,Fast Text和TextCNN模型分别在4种类型和36种亚型的意图分类中取得了最好的结果。结论在这项工作中,我们提供了一个用于中国医疗意图分类的数据集,可用于医疗质量保证及相关领域。我们在CMID上执行了一个意图分类任务。此外,我们还对数据集的内容进行了分析。
BackgroundTo provide satisfying answers, medical QA system has to understand the intentions of the users' questions precisely. For medical intent classification, it requires high-quality datasets to train a deep-learning approach in a supervised way. Currently, there is no public dataset for Chinese medical intent classification, and the datasets of other fields are not applicable to the medical QA system. To solve this problem, we construct a Chinese medical intent dataset (CMID) using the questions from medical QA websites. On this basis, we compare four intent classification models on CMID using a case study.MethodsThe questions in CMID are obtained from several medical QA websites. The intent annotation standard is developed by the medical experts, which includes four types and 36 subtypes of users' intents. Besides the intent label, CMID also provides two types of additional information, including word segmentation and named entity. We use the crowdsourcing way to annotate the intent information for each Chinese medical question. Word segmentation and named entities are obtained using the Jieba and a well-trained Lattice-LSTM model. We loaded a Chinese medical dictionary consisting of 530,000 for word segmentation to obtain a more accurate result. We also select four popular deep learning-based models and compare their performances of intent classification on CMID.ResultsThe final CMID contains 12,000 Chinese medical questions and is organized in JSON format. Each question is labeled the intention, word segmentation, and named entity information. The information about question length, number of entities, and are also detailed analyzed. Among Fast Text, TextCNN, TextRNN, and TextGCN, Fast Text and TextCNN models have achieved the best results in four types and 36 subtypes intent classification, respectively.ConclusionsIn this work, we provide a dataset for Chinese medical intent classification, which can be used in medical QA and related fields. We performed an intent classification task on the CMID. In addition, we also did some analysis on the content of the dataset.