Multigranularity Label Prediction Model for Automatic International Classification of Diseases Coding in Clinical Text

Multigranularity Label Prediction Model for Automatic International Classification of Diseases Coding in Clinical Text
复制标题

DOI:
10.1089/cmb.2023.0096
复制
发表时间:
2023-07
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Ying Yu;Tian Qiu;Junwen Duan;Jianxin Wang
Ying Yu;Tian Qiu;Junwen Duan;Jianxin Wang
中科院分区:
其他
文献类型:
--
作者:
Ying Yu;Tian Qiu;Junwen Duan;Jianxin Wang

文献摘要

相似文献

《国际疾病分类》是产生跨区域和跨时间的可比较全球疾病统计数据的基础。ICD编码的过程包括根据临床记录为疾病分配代码,临床记录可以以标准的方式描述患者的病情。然而,这一过程由于大量的代码和ICD代码的复杂分类而变得复杂,这些代码按层次组织为不同的级别,包括章节、类别、子类别及其细分。现有的许多研究只关注子类别码的预测,忽略了码之间的层次关系。为了解决这一限制,我们提出了一个多任务学习模型,该模型为不同的代码级别训练多个分类器,同时还通过强化机制捕获粗粒度和细粒度标签之间的关系。我们的方法在英文和中文基准数据集上进行了评估,我们证明了我们的方法在基线模型上取得了具有竞争力的表现,特别是在宏观f1结果方面。这些发现表明,我们的方法有效地利用了ICD代码的层次结构来提高疾病代码预测的准确性。注意机制分析表明,该模型的多粒度注意在不同的粒度层次上捕捉了输入文本的关键特征,可以为预测结果提供合理的解释。
International Classification of Diseases (ICD) serves as the foundation for generating comparable global disease statistics across regions and over time. The process of ICD coding involves assigning codes to diseases based on clinical notes, which can describe a patient's condition in a standard way. However, this process is complicated by the vast number of codes and the intricate taxonomy of ICD codes, which are hierarchically organized into various levels, including chapter, category, subcategory, and its subdivisions. Many existing studies focus solely on predicting subcategory codes, ignoring the hierarchical relationships among codes. To address this limitation, we propose a multitask learning model that trains multiple classifiers for different code levels, while also capturing the relations between coarser and finer-grained labels through a reinforcement mechanism. Our approach is evaluated on both English and Chinese benchmark dataset, and we demonstrate that our method achieves competitive performance with baseline models, particularly in terms of macro-F1 results. These findings suggest that our approach effectively leverages the hierarchical structure of ICD codes to improve disease code prediction accuracy. Analysis of attention mechanism shows that multigranularity attention of our model captures crucial feature of input text on different granularity levels, which can provide reasonable explanations for the prediction results.