Classification of Current Procedural Terminology Codes from Electronic Health Record Data Using Machine Learning.

Classification of Current Procedural Terminology Codes from Electronic Health Record Data Using Machine Learning.
复制标题

DOI:
10.1097/aln.0000000000003150
复制
发表时间:
2020-04
期刊:
影响因子:
8.8
通讯作者:
Saager L
Saager L
中科院分区:
医学1区
文献类型:
--
作者:
Burns ML;Mathis MR;Vandervest J;Tan X;Lu B;Colquhoun DA;Shah N;Kheterpal S;Saager L

文献摘要

被引文献

相似文献

准确的麻醉程序代码数据对于麻醉实践中的质量改进、研究和报销任务至关重要。先进的数据科学技术,包括机器学习和自然语言处理,为麻醉程序中的当前程序术语代码开发分类工具提供了机会。使用Train/Test数据集创建模型,该数据集包括来自16家学术和私立医院的1,164,343个程序。创建了五个监督机器学习模型,用于对麻醉学当前手术术语代码进行分类,准确性定义为与围手术期数据库中现有的机构分配代码匹配的首选分类。这两个表现最好的模型被进一步细化,并在来自与Train/Test不同的单一机构的Holdout数据集上进行测试。创建了一个可调置信参数,以识别模型高度准确的病例,目标是≥ 95%的准确度,高于2018年医疗保险和医疗补助服务中心报告的按服务收费准确度。实际提交的索赔数据从计费专家被用作参考标准。支持向量机和神经网络标签嵌入注意模型是表现最好的模型,分别展示了87.9%和84.2%(单一最佳代码)以及96.8%和94.0%(前三名)的整体准确率。在Train/Test数据集中,使用支持向量机的47.0%的案例分类准确率为96.4%,使用标签嵌入注意模型的62.2%的案例分类准确率为94.4%。在Holdout数据集中,58.0%病例的分类准确率为93.1%,62.0%病例的分类准确率为95.0%。模型训练中最重要的特征是过程文本。通过应用机器学习和自然语言处理技术,为麻醉学当前程序术语代码分类创建了高度准确的实时模型。这种分类方法的增加的处理速度和先验目标准确性可以为依赖于麻醉程序代码的质量改进、研究和报销任务提供性能优化和成本降低。
Accurate anesthesiology procedure code data is essential to quality improvement, research, and reimbursement tasks within anesthesiology practices. Advanced data science techniques including machine learning and natural language processing offer opportunities to develop classification tools for Current Procedural Terminology codes across anesthesia procedures. Models were created using a Train/Test dataset including 1,164,343 procedures from 16 academic and private hospitals. Five supervised machine learning models were created to classify anesthesiology Current Procedural Terminology codes, with accuracy defined as first choice classification matching the institutional-assigned code existing in the perioperative database. The two best performing models were further refined and tested on a Holdout dataset from a single institution distinct from Train/Test. A tunable confidence parameter was created to identify cases for which models were highly accurate, with the goal of ≥95% accuracy, above the reported 2018 Centers for Medicare and Medicaid Services fee-for-service accuracy. Actual submitted claim data from billing specialists was used as a reference standard. Support vector machine and neural network label-embedding attentive models were the best performing models, respectively demonstrating overall accuracies of 87.9% and 84.2% (single best code), and 96.8% and 94.0% (within top three). Classification accuracy was 96.4% in 47.0% of cases using support vector machine and 94.4% in 62.2% of cases using label-embedding attentive model within the Train/Test dataset. In the Holdout dataset, respective classification accuracies were 93.1% in 58.0% of cases and 95.0% among 62.0%. The most important feature in model training was procedure text. Through application of machine learning and natural language processing techniques, highly accurate real-time models were created for anesthesiology Current Procedural Terminology code classification. The increased processing speed and a priori targeted accuracy of this classification approach may provide performance optimization and cost reduction for quality improvement, research, and reimbursement tasks reliant on anesthesiology procedure codes.