Development of a Portable Tool to Identify Patients With Atrial Fibrillation Using Clinical Notes From the Electronic Medical Record.

Development of a Portable Tool to Identify Patients With Atrial Fibrillation Using Clinical Notes From the Electronic Medical Record.
复制标题

DOI:
10.1161/circoutcomes.120.006516
复制
发表时间:
2020-10
期刊:
Circulation. Cardiovascular quality and outcomes
影响因子:
--
通讯作者:
Lloyd-Jones DM
Lloyd-Jones DM
中科院分区:
其他
文献类型:
--
作者:
Shah RU;Mutharasan RK;Ahmad FS;Rosenblatt AG;Gay HC;Steinberg BA;Yandell M;Tristani-Firouzi M;Klewer J;Mukherjee R;Lloyd-Jones DM

文献摘要

相似文献

电子病历包含大量隐藏在自由文本中的信息。我们创建了一种自然语言处理 (NLP) 算法,仅使用文本即可识别心房颤动 (AF) 患者。我们从 2010 年至 2017 年期间至少拥有一个 AF 计费代码的患者创建了三个数据集:训练集 (n=886)、来自站点 #1 的内部验证集 (n=285) 和来自站点 #2 的外部验证集 (n=276)。一组临床医生对患者进行审查并判定是否存在房颤,并以此作为参考标准。我们训练了 54 种算法来对每位患者进行分类,改变模型、特征数量、停用词数量以及用于创建特征集的方法。将训练集中具有最高 F 分数(灵敏度和阳性预测值的调和平均值)的算法应用于验证集。使用自举法比较站点 #1 和站点 #2 之间的 F 分数和接收者操作特征曲线下面积 (AUC)。判定的 1 号站点 AF 患病率为 75.1%,2 号站点为 86.2%。在 54 种算法中,表现最好的模型是逻辑回归,使用 1000 个特征、100 个停用词和词频-逆文档频率(TF-IDF)方法创建特征集,训练集中的灵敏度为 92.8%,特异性为 93.9%,AUC 为 0.93。位点#1 的敏感性为 92.5%,特异性为 88.7%,AUC 为 0.91。位点#2 的表现是敏感性 89.5%,特异性 71.1%,AUC 为 0.80。与站点 #1 相比,站点 #2 的 F 分数较低(92.5% [SD 1.1%] 对比 94.2% [SD 1.1%];p<0.001)。我们开发了一种 NLP 算法,仅使用文本即可识别 AF 患者,在两个不同站点的 F 分数均超过 90%。这种方法可以更好地利用临床叙述,并为精确、高通量的队列识别创造机会。
The electronic medical record contains a wealth of information buried in free text. We created a natural language processing (NLP) algorithm to identify atrial fibrillation (AF) patients using text alone. We created three data sets from patients with at least one AF billing code from 2010 to 2017: a training set (n=886), an internal validation set from Site #1 (n=285), and an external validation set from Site #2 (n=276). A team of clinicians reviewed and adjudicated patients as AF present or absent, which served as the reference standard. We trained 54 algorithms to classify each patient, varying the model, number of features, number of stop words, and the method used to create the feature set. The algorithm with the highest F-score (the harmonic mean of sensitivity and positive predictive value) in the training set was applied to the validation sets. F-scores and area under the receiver operating characteristic curves (AUC) were compared between Site #1 and Site #2 using bootstrapping. Adjudicated AF prevalence was 75.1% at Site #1 and 86.2% at Site #2. Among 54 algorithms, the best performing model was logistic regression, using 1000 features, 100 stop words, and term frequency-inverse document frequency (TF-IDF) method to create the feature set, with sensitivity 92.8%, specificity 93.9%, and an AUC of 0.93 in the training set. The performance at Site #1 was sensitivity 92.5%, specificity 88.7%, with an AUC of 0.91. The performance at Site #2 was sensitivity 89.5%, specificity 71.1%, with an AUC of 0.80. The F-score was lower at Site #2 compared to Site #1 (92.5% [SD 1.1%] versus 94.2% [SD 1.1%]; p<0.001). We developed a NLP algorithm to identify AF patients using text alone, with >90% F-score at two separate sites. This approach allows better use of the clinical narrative, and creates an opportunity for precise, high throughput cohort identification.