Text mining brain imaging reports

Text mining brain imaging reports
复制标题

DOI:
10.1186/s13326-019-0211-7
复制
发表时间:
2019-11-12
影响因子:
1.9
通讯作者:
Whiteley, William
Whiteley, William
中科院分区:
工程技术4区
文献类型:
--
作者:
Alex, Beatrice;Grover, Claire;Whiteley, William

文献摘要

被引文献

相似文献

背景随着文本挖掘技术的改进和大型非结构化电子医疗记录(EHR)数据集的可用性,现在可以以相当高的精度从EHR中包含的原始文本中提取结构化信息。我们描述了一个文本挖掘系统,用于对放射科医生的CT和MRI脑部扫描报告进行分类,分配指示中风发生和类型以及其他观察结果的标签。我们的系统,爱丁堡信息提取放射学报告(EdIE-R)系统,我们在这里描述,开发和测试的放射学reports.The工作在本文中报告的集合是基于1168放射学报告从爱丁堡中风研究(ESS),一个医院为基础的登记中风和短暂性脑缺血发作患者。我们为这些数据手动创建注释,同时开发基于规则的EdIE-R系统,以识别放射学报告中与卒中相关的表型信息。这个过程是迭代的,每次迭代都考虑领域专家的反馈,以适应和调整EdIE-R文本挖掘系统,该系统识别每个报告中的实体,否定和实体之间的关系,并确定报告级别的标签(表型)。结果所有类型标注的标注者间一致性(IAA)较高,实体标注为96.96,否定标注为96.46,关系标注为95.84,标签标注为94.02。盲测试集上的等效系统分数同样高,对于第一个注释器,实体为95.49,否定为94.41,关系为98.27,标签为96.39,对于第二个注释器,分别为96.86,96.01,96.53和92.61。结论以如此高的准确度自动阅读此类EHR数据为人群健康监测和审计开辟了途径,并可为流行病学研究提供资源。我们正在英格兰和苏格兰的NHS中验证EdIE-R。人工标注的ESS语料库将可用于应用研究目的。
Background With the improvements to text mining technology and the availability of large unstructured Electronic Healthcare Records (EHR) datasets, it is now possible to extract structured information from raw text contained within EHR at reasonably high accuracy. We describe a text mining system for classifying radiologists' reports of CT and MRI brain scans, assigning labels indicating occurrence and type of stroke, as well as other observations. Our system, the Edinburgh Information Extraction for Radiology reports (EdIE-R) system, which we describe here, was developed and tested on a collection of radiology reports.The work reported in this paper is based on 1168 radiology reports from the Edinburgh Stroke Study (ESS), a hospital-based register of stroke and transient ischaemic attack patients. We manually created annotations for this data in parallel with developing the rule-based EdIE-R system to identify phenotype information related to stroke in radiology reports. This process was iterative and domain expert feedback was considered at each iteration to adapt and tune the EdIE-R text mining system which identifies entities, negation and relations between entities in each report and determines report-level labels (phenotypes). Results The inter-annotator agreement (IAA) for all types of annotations is high at 96.96 for entities, 96.46 for negation, 95.84 for relations and 94.02 for labels. The equivalent system scores on the blind test set are equally high at 95.49 for entities, 94.41 for negation, 98.27 for relations and 96.39 for labels for the first annotator and 96.86, 96.01, 96.53 and 92.61, respectively for the second annotator. Conclusion Automated reading of such EHR data at such high levels of accuracies opens up avenues for population health monitoring and audit, and can provide a resource for epidemiological studies. We are in the process of validating EdIE-R in separate larger cohorts in NHS England and Scotland. The manually annotated ESS corpus will be available for research purposes on application.