ICD-10 code retrieval based on distributional semantics of diagnosis descriptions

ICD-10 code retrieval based on distributional semantics of diagnosis descriptions
复制标题

DOI:
10.1109/icaicta.2017.8090957
复制
发表时间:
2017-08
期刊:
2017 International Conference on Advanced Informatics, Concepts, Theory, and Applications (ICAICTA)
影响因子:
--
通讯作者:
T. Akiba;B. Sy;Ayman Zeidan
T. Akiba;B. Sy;Ayman Zeidan
中科院分区:
其他
文献类型:
--
作者:
T. Akiba;B. Sy;Ayman Zeidan

文献摘要

相似文献

在本文中,我们提出了一种从患者疾病投诉的自然语言描述中提取ICD-10代码的方法。所提出的方法是基于出现在两种自然语言表达中的术语的分布语义:患者的抱怨和ICD-10代码描述。为了在给定的长而嘈杂的患者表达中定位相关的单词片段,在评估患者的投诉与ICD代码之间的匹配之前,进行单词对单词的比对。用于初步研究的数据集包括81例测试患者记录。对于每条记录,拟议的系统从总共69,000个ICD-10代码中检索一组代码。通过实验评估,我们发现,平均而言,系统能够从前10个结果中返回3.6个正确的代码。通过利用用户的交互,性能得到了进一步提高,从前10个代码中建议了大约4个正确的代码。
In this paper, we propose a method for extracting ICD-10 codes from the natural language description of a patient illness complaint. The proposed method is based on distributional semantics of terms that appeared in the two natural language expressions: a patient's complaint and an ICD-10 code description. In order to locate the relevant fragment of words within a given long and noisy patient's expression, word-to-word alignment is performed before evaluating the match between a patient's complaint and an ICD code. The data set used for the preliminary study consists of 81 test patient records. For each record, the proposed system retrieves a set of codes from a total of 69,000 ICD-10 codes. Through the experimental evaluation, we found that, on average, the system was able to return 3.6 correct codes from the top 10 results. By making use of a user's interaction, the performance was further improved to suggest about four correct codes from the top 10.