Automatic classification of scanned electronic health record documents.

Automatic classification of scanned electronic health record documents.
复制标题

DOI:
10.1016/j.ijmedinf.2020.104302
复制
发表时间:
2020-12
影响因子:
4.9
通讯作者:
Bernstam EV
Bernstam EV
中科院分区:
医学2区
文献类型:
--
作者:
Goodrum H;Roberts K;Bernstam EV

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)包含来自各种来源的扫描文档,例如身份证、放射学报告、临床信函和许多其他文档类型。我们描述了在一个卫生机构的扫描文件的分布,并描述了一个系统的设计和评估,以分类文件到临床相关和非临床相关的类别,以及进一步的子分类。我们的目标是证明文本分类系统可以准确地对扫描文档进行分类。我们使用光学字符识别(OCR)提取文本。然后,我们创建并评估了多个文本分类机器学习模型,包括“词袋”和深度学习方法。我们使用整个文档作为输入,以及文档的各个页面,在三个不同的分类级别上对系统进行了评估。最后,我们比较了不同的文本处理方法的效果。使用ClinicalBERT的深度学习模型表现最好。该模型以0.973的准确度区分临床相关文档和非临床相关文档;以0.949的准确度区分中间子分类;以0.913的准确度区分单个类别。在EHR中,某些文档类别(如“外部医疗记录”)可能包含数百个扫描页面,而没有明确的文档边界。如果没有进一步的分类,临床医生必须查看每一页,否则可能会丢失临床相关信息。机器学习可以自动对这些扫描文档进行分类,以减轻临床医生的负担。将机器学习应用于OCR提取的文本有可能准确识别EHR中与临床相关的扫描内容。
Electronic Health Records (EHRs) contain scanned documents from a variety of sources such as identification cards, radiology reports, clinical correspondence, and many other document types. We describe the distribution of scanned documents at one health institution and describe the design and evaluation of a system to categorize documents into clinically relevant and non-clinically relevant categories as well as further sub-classifications. Our objective is to demonstrate that text classification systems can accurately classify scanned documents. We extracted text using Optical Character Recognition (OCR). We then created and evaluated multiple text classification machine learning models, including both “bag of words” and deep learning approaches. We evaluated the system on three different levels of classification using both the entire document as input, as well as the individual pages of the document. Finally, we compared the effects of different text processing methods. A deep learning model using ClinicalBERT performed best. This model distinguished between clinically-relevant documents and not clinically-relevant documents with an accuracy of 0.973; between intermediate sub-classifications with an accuracy of 0.949; and between individual classes with an accuracy of 0.913. Within the EHR, some document categories such as “external medical records” may contain hundreds of scanned pages without clear document boundaries. Without further sub-classification, clinicians must view every page or risk missing clinically-relevant information. Machine learning can automatically classify these scanned documents to reduce clinician burden. Using machine learning applied to OCR-extracted text has the potential to accurately identify clinically-relevant scanned content within EHRs.
DOI: 10.1186/s13326-017-0120-6
发表时间: 2017-03-03
影响因子: 1.9
作者:
Du J;Xu J;Song H;Liu X;Tao C
通讯作者: Tao C
DOI: 10.1197/jamia.m1474
发表时间: 2004-05-01
影响因子: 6.4
作者:
Cowell, J;Zeng, Q;Lacroix, EM
通讯作者: Lacroix, EM
DOI: 10.1038/nature21056
发表时间: 2017-02-02
期刊: Nature
影响因子: 64.8
作者:
Esteva A;Kuprel B;Novoa RA;Ko J;Swetter SM;Blau HM;Thrun S
通讯作者: Thrun S
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.1136/amiajnl-2013-001686
发表时间: 2014-02-01
影响因子: 6.4
作者:
Friedman, Asia;Crosson, Jesse C.;Cohen, Deborah J.
通讯作者: Cohen, Deborah J.