Feature engineering from medical notes: A case study of dementia detection.

Feature engineering from medical notes: A case study of dementia detection.
复制标题

医学笔记的特征工程:痴呆症检测的案例研究。

DOI:
10.1016/j.heliyon.2023.e14636
复制
发表时间:
2023
期刊:
影响因子:
4
通讯作者:
Boustani,Malaz
Boustani,Malaz
中科院分区:
综合性期刊4区
文献类型:
--
作者:
BenMiled,Zina;Dexter,PaulR;Grout,RandallW;Boustani,Malaz

文献摘要

被引文献

相似文献

背景和目的医学笔记是以自由文本格式描述病人健康状况的叙述。这些记录可以比结构化数据(如药物史或疾病状况)提供更多信息。这些数据是常规收集的,可用于评估患者患痴呆症等慢性疾病的风险。本研究探讨了将常规护理笔记转化为痴呆风险分类器的不同方法,并评估了这些分类器对新患者和新医疗机构的普遍性。方法对患者的相关病史进行详细的记录。在本研究中,使用TF-ICF选择高危痴呆患者与健康对照之间判别能力最高的关键词。然后以所选关键字的出现次数形式总结医疗记录。比较了两种不同的摘要编码。第一个编码由BERT或临床BERT预训练语言模型产生的每个关键字出现的向量嵌入的平均值组成。第二种编码方式根据UMLS概念聚合关键字,并使用每个概念作为曝光变量。对于这两种编码,还考虑了所选关键字的拼写错误,以提高分类器的预测性能。在第一次编码上建立了神经网络,在第二次编码上应用了梯度增强树模型。使用来自单一医疗保健机构的患者来开发所有分类器,然后对来自同一医疗保健机构的滞留患者以及来自其他两个医疗保健机构的测试患者进行评估。结果表明,当梯度增强树模型与来自UMLS概念的暴露变量结合使用时,使用AUC为75%的医疗记录,可以在发病前一年识别出有痴呆风险的患者。然而,当嵌入特征空间和分类器应用于其他医疗机构的患者时,这种性能无法保持。此外,对梯度增强树模型的顶级预测因子的分析表明,根据是否包含关键字的拼写变体,不同的特征通知分类。结论医学笔记可以建立痴呆等复杂慢性疾病的风险预测模型。然而,需要进一步的研究工作来提高这些模型的可泛化性。这些努力应考虑到医疗记录的长度和地点;为每一种疾病提供足够的训练数据;以及不同的特征工程技术所带来的可变性。
Background and objectivesMedical notes are narratives that describe the health of the patient in free text format. These notes can be more informative than structured data such as the history of medications or disease conditions. They are routinely collected and can be used to evaluate the patient's risk for developing chronic diseases such as dementia. This study investigates different methodologies for transforming routine care notes into dementia risk classifiers and evaluates the generalizability of these classifiers to new patients and new health care institutions.MethodsThe notes collected over the relevant history of the patient are lengthy. In this study, TF-ICF is used to select keywords with the highest discriminative ability between at risk dementia patients and healthy controls. The medical notes are then summarized in the form of occurrences of the selected keywords. Two different encodings of the summary are compared. The first encoding consists of the average of the vector embedding of each keyword occurrence as produced by the BERT or Clinical BERT pre-trained language models. The second encoding aggregates the keywords according to UMLS concepts and uses each concept as an exposure variable. For both encodings, misspellings of the selected keywords are also considered in an effort to improve the predictive performance of the classifiers. A neural network is developed over the first encoding and a gradient boosted trees model is applied to the second encoding. Patients from a single health care institution are used to develop all the classifiers which are then evaluated on held-out patients from the same health care institution as well as test patients from two other health care institutions.ResultsThe results indicate that it is possible to identify patients at risk for dementia one year ahead of the onset of the disease using medical notes with an AUC of 75% when a gradient boosted trees model is used in conjunction with exposure variables derived from UMLS concepts. However, this performance is not maintained with an embedded feature space and when the classifier is applied to patients from other health care institutions. Moreover, an analysis of the top predictors of the gradient boosted trees model indicates that different features inform the classification depending on whether or not spelling variants of the keywords are included.ConclusionThe present study demonstrates that medical notes can enable risk prediction models for complex chronic diseases such as dementia. However, additional research efforts are needed to improve the generalizability of these models. These efforts should take into consideration the length and localization of the medical notes; the availability of sufficient training data for each disease condition; and the variabilities resulting from different feature engineering techniques.