A validated natural language processing algorithm for brain imaging phenotypes from radiology reports in UK electronic health records

A validated natural language processing algorithm for brain imaging phenotypes from radiology reports in UK electronic health records
复制标题

DOI:
10.1186/s12911-019-0908-7
复制
发表时间:
2019-09-09
影响因子:
3.5
通讯作者:
Whiteley, William
Whiteley, William
中科院分区:
医学3区
文献类型:
--
作者:
Wheater, Emily;Mair, Grant;Whiteley, William

文献摘要

被引文献

相似文献

脑放射学报告中表型的手工编码是耗时的。我们开发了一种自然语言处理(NLP)算法,用于自动识别英国国家卫生服务(NHS)常规临床实践中放射学报告中的脑成像。方法:我们使用来自卒中/TIA患者队列研究和地区医院的匿名文本脑成像报告来开发和测试NLP算法。两位专家在1692年的报告中标记了24种脑血管和其他神经系统表型的文本。我们首先在队列研究中开发并测试了基于规则的NLP算法,并在地区医院的报告中进一步对其进行了评估。结果在两个数据集上,专家读者之间的一致性非常好(Cohen’s kappa =0.93)。在未见过的地区医院报告的最终测试数据集(n = 700)中,该算法对任何缺血性卒中的报告都有非常好的表现[灵敏度89% (95% CI:81-94);阳性预测值(PPV) 85% (76 ~ 90);特异性100% (95% CI:0.99-1.00);任何出血性卒中[敏感性96% (95% CI: 80-99), PPV 72% (95% CI:55-84);特异性100% (95% CI:0.99-1.00);脑肿瘤[敏感性96% (CI:87-99);PPV 84% (73-91);特异性:100% (95% CI:0.99-1.00)]和脑小血管疾病和脑萎缩(敏感性、PPV和特异性均为97%)。我们获得少数蛛网膜下腔出血、微出血或硬膜下血肿的报告。在来自NHS Tayside的110,695份报告中,萎缩(n = 28,757, 26%)、小血管疾病(15,015,14%)和老年深部缺血性中风(10,636,10%)是最常见的发现。NLP算法可以在英国NHS放射学记录中开发,以允许在规模上识别具有重要脑成像表型的患者队列,否则这是不可能的。
Background Manual coding of phenotypes in brain radiology reports is time consuming. We developed a natural language processing (NLP) algorithm to enable automatic identification of brain imaging in radiology reports performed in routine clinical practice in the UK National Health Service (NHS). Methods We used anonymized text brain imaging reports from a cohort study of stroke/TIA patients and from a regional hospital to develop and test an NLP algorithm. Two experts marked up text in 1692 reports for 24 cerebrovascular and other neurological phenotypes. We developed and tested a rule-based NLP algorithm first within the cohort study, and further evaluated it in the reports from the regional hospital. Results The agreement between expert readers was excellent (Cohen's kappa =0.93) in both datasets. In the final test dataset (n = 700) in unseen regional hospital reports, the algorithm had very good performance for a report of any ischaemic stroke [sensitivity 89% (95% CI:81-94); positive predictive value (PPV) 85% (76-90); specificity 100% (95% CI:0.99-1.00)]; any haemorrhagic stroke [sensitivity 96% (95% CI: 80-99), PPV 72% (95% CI:55-84); specificity 100% (95% CI:0.99-1.00)]; brain tumours [sensitivity 96% (CI:87-99); PPV 84% (73-91); specificity: 100% (95% CI:0.99-1.00)] and cerebral small vessel disease and cerebral atrophy (sensitivity, PPV and specificity all > 97%). We obtained few reports of subarachnoid haemorrhage, microbleeds or subdural haematomas. In 110,695 reports from NHS Tayside, atrophy (n = 28,757, 26%), small vessel disease (15,015, 14%) and old, deep ischaemic strokes (10,636, 10%) were the commonest findings. Conclusions An NLP algorithm can be developed in UK NHS radiology records to allow identification of cohorts of patients with important brain imaging phenotypes at a scale that would otherwise not be possible.