A numerical similarity approach for using retired Current Procedural Terminology (CPT) codes for electronic phenotyping in the Scalable Collaborative Infrastructure for a Learning Health System (SCILHS).

A numerical similarity approach for using retired Current Procedural Terminology (CPT) codes for electronic phenotyping in the Scalable Collaborative Infrastructure for a Learning Health System (SCILHS).
复制标题

DOI:
10.1186/s12911-015-0223-x
复制
发表时间:
2015-12-11
影响因子:
3.5
通讯作者:
Murphy SN
Murphy SN
中科院分区:
医学3区
文献类型:
--
作者:
Klann JG;Phillips LC;Turchin A;Weiler S;Mandl KD;Murphy SN

文献摘要

被引文献

相似文献

识别符合观察性研究或临床试验合格性标准的患者队列所需的互操作表型分析算法需要一致的结构化编码格式的医学数据。数据异构性限制了这些算法的适用性。现有的方法通常是:没有广泛的互操作性;或者,由于依赖于最低公分母(ICD-9诊断)而具有低灵敏度。在学习型医疗保健系统的可扩展协作基础设施(SCILHS)中,我们奋进使用广泛可用的当前程序术语(CPT)程序代码和ICD-9。不幸的是,CPT每年都有很大的变化-代码被淘汰/替换。纵向分析需要对已停用和当前代码进行分组。BioPortal提供了一个可导航的CPT层次结构,我们将其导入到Informatics for Integrating Biology and the Bedside(i2 b2)数据仓库和分析平台中。但是,此层次结构不包括失效代码。我们将BioPortal的2014 AA CPT层次结构与Partners Healthcare的SCILHS数据集市进行了比较,其中包括15年来300万患者的数据。2014 AA中有573个CPT代码不存在(650万次)。没有现有的术语提供这些缺失的代码的层次联系,所以我们开发了一种方法,自动将缺失的代码在最具体的“石斑鱼”类别,使用CPT代码的数值相似性。两名信息学家审查了结果。我们将最终的表纳入我们的i2 b2 SCILHS/PCORnet本体中,在七个站点部署它,并对几种表型分析算法进行了差距分析和评估。评审人员发现,当只考虑错误分类时,该方法以97%的精度正确放置代码(“正确性精度”),使用最佳放置的黄金标准时,该方法的精度为52%(“最优精度”)。高正确性精度意味着代码被放置在一个合理的层次位置,审查人员可以快速验证。较低的最优性精度意味着代码不经常被放置在最优层次子文件夹中。这七个网站很少遇到我们本体之外的代码,其中93%只包含四个代码。我们的分层方法正确地分组退休和非退休代码在大多数情况下,并延长了几个重要的表型分析算法的时间范围。我们开发了一种简单,易于验证,自动化的方法,将退休的CPT代码放入BioPortal CPT层次结构。这补充了现有的分层术语,其中不包括退休的代码。该方法的实用性被证实的高正确性精度和成功的退休与非退休代码的分组。
Interoperable phenotyping algorithms, needed to identify patient cohorts meeting eligibility criteria for observational studies or clinical trials, require medical data in a consistent structured, coded format. Data heterogeneity limits such algorithms’ applicability. Existing approaches are often: not widely interoperable; or, have low sensitivity due to reliance on the lowest common denominator (ICD-9 diagnoses). In the Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS) we endeavor to use the widely-available Current Procedural Terminology (CPT) procedure codes with ICD-9. Unfortunately, CPT changes drastically year-to-year – codes are retired/replaced. Longitudinal analysis requires grouping retired and current codes. BioPortal provides a navigable CPT hierarchy, which we imported into the Informatics for Integrating Biology and the Bedside (i2b2) data warehouse and analytics platform. However, this hierarchy does not include retired codes. We compared BioPortal’s 2014AA CPT hierarchy with Partners Healthcare’s SCILHS datamart, comprising three-million patients’ data over 15 years. 573 CPT codes were not present in 2014AA (6.5 million occurrences). No existing terminology provided hierarchical linkages for these missing codes, so we developed a method that automatically places missing codes in the most specific “grouper” category, using the numerical similarity of CPT codes. Two informaticians reviewed the results. We incorporated the final table into our i2b2 SCILHS/PCORnet ontology, deployed it at seven sites, and performed a gap analysis and an evaluation against several phenotyping algorithms. The reviewers found the method placed the code correctly with 97 % precision when considering only miscategorizations (“correctness precision”) and 52 % precision using a gold-standard of optimal placement (“optimality precision”). High correctness precision meant that codes were placed in a reasonable hierarchal position that a reviewer can quickly validate. Lower optimality precision meant that codes were not often placed in the optimal hierarchical subfolder. The seven sites encountered few occurrences of codes outside our ontology, 93 % of which comprised just four codes. Our hierarchical approach correctly grouped retired and non-retired codes in most cases and extended the temporal reach of several important phenotyping algorithms. We developed a simple, easily-validated, automated method to place retired CPT codes into the BioPortal CPT hierarchy. This complements existing hierarchical terminologies, which do not include retired codes. The approach’s utility is confirmed by the high correctness precision and successful grouping of retired with non-retired codes.