Importance of multi-modal approaches to effectively identify cataract cases from electronic health records

Importance of multi-modal approaches to effectively identify cataract cases from electronic health records
复制标题

DOI:
10.1136/amiajnl-2011-000456
复制
发表时间:
2012-03-01
影响因子:
6.4
通讯作者:
Starren, Justin B.
Starren, Justin B.
中科院分区:
管理学2区
文献类型:
--
作者:
Peissig, Peggy L.;Rasmussen, Luke V.;Starren, Justin B.

文献摘要

被引文献

相似文献

目的利用电子健康记录(EHR)识别受试者进行基因组关联研究的兴趣越来越大,部分原因是大量临床数据的可用性和受试者识别的预期成本效益。我们描述了一个EHR为基础的算法,以确定受试者与年龄相关的cataracts.Materials和方法的建设和验证,我们使用了多模式的策略,包括结构化的数据库查询,自然语言处理的自由文本文件,光学字符识别扫描的临床图像,以确定白内障科目和相关的白内障属性。对3657名受试者进行了广泛的验证,将多模态结果与手动图表审查进行了比较。该算法也在参与的电子医学记录和基因组学(eMERGE.Results)institutionals.Results基于EHR的白内障表型分型算法成功开发和验证,导致阳性预测值(PPV)> 95%。与单模式方法相比,多模式方法将白内障受试者属性的识别提高了三倍,同时保持了高PPV。组件的白内障算法成功地部署在其他三个机构具有类似的accuracy.Discussion一个多模式的战略,将光学字符识别和自然语言处理可能会增加的情况下,同时保持类似的PPV确定的数量。然而,这样的算法,需要将所需的信息嵌入到临床documents.Conclusion我们已经证明,算法来识别和表征白内障可以开发利用收集的数据通过EHR。这些算法即使在跨多个EHR和机构边界实施时也能提供高水平的准确性。
Objective There is increasing interest in using electronic health records (EHRs) to identify subjects for genomic association studies, due in part to the availability of large amounts of clinical data and the expected cost efficiencies of subject identification. We describe the construction and validation of an EHR-based algorithm to identify subjects with age-related cataracts.Materials and methods We used a multi-modal strategy consisting of structured database querying, natural language processing on free-text documents, and optical character recognition on scanned clinical images to identify cataract subjects and related cataract attributes. Extensive validation on 3657 subjects compared the multi-modal results to manual chart review. The algorithm was also implemented at participating electronic MEdical Records and GEnomics (eMERGE) institutions.Results An EHR-based cataract phenotyping algorithm was successfully developed and validated, resulting in positive predictive values (PPVs) >95%. The multi-modal approach increased the identification of cataract subject attributes by a factor of three compared to single-mode approaches while maintaining high PPV. Components of the cataract algorithm were successfully deployed at three other institutions with similar accuracy.Discussion A multi-modal strategy incorporating optical character recognition and natural language processing may increase the number of cases identified while maintaining similar PPVs. Such algorithms, however, require that the needed information be embedded within clinical documents.Conclusion We have demonstrated that algorithms to identify and characterize cataracts can be developed utilizing data collected via the EHR. These algorithms provide a high level of accuracy even when implemented across multiple EHRs and institutional boundaries.