Discretization of continuous features in clinical datasets

Discretization of continuous features in clinical datasets
复制标题

DOI:
10.1136/amiajnl-2012-000929
复制
发表时间:
2013-05-01
影响因子:
6.4
通讯作者:
Lowe, Henry J.
Lowe, Henry J.
中科院分区:
管理学2区
文献类型:
--
作者:
Maslove, David M.;Podchiyska, Tanya;Lowe, Henry J.

文献摘要

被引文献

相似文献

背景来自电子病历(EMR)的临床数据的日益可获得性为健康信息的二次利用创造了机会。目的利用EMR数据对6种离散化策略进行评价。材料与方法利用决策树和朴素贝叶斯分类器对重症监护病房成年患者的实验室数据(动脉血气)和生理数据(心输出量)进行分类。使用两种监督离散化策略和四种非监督离散化策略对连续特征进行分割。结果有监督的方法比无监督的方法更准确、更一致,但往往会产生更大的决策树。在非监督方法中,等频率和k-均值总体表现较好,而等宽度方法的准确率明显较低。讨论我们相信,这是第一次使用电子病历数据对离散化策略进行专门评估。任何一种离散化方法不太可能普遍适用于电子病历数据。性能受类别标签的选择以及在非监督方法的情况下的间隔数目的影响。在选择区间数时,通常在更高的准确性和更大的一致性之间进行权衡。结论总的来说,有监督的方法可以产生更高的精度,但仅限于单个特定的应用。非监督方法不需要类标签,并且可以产生可用于多种目的的离散化数据。
Background The increasing availability of clinical data from electronic medical records (EMRs) has created opportunities for secondary uses of health information. When used in machine learning classification, many data features must first be transformed by discretization.Objective To evaluate six discretization strategies, both supervised and unsupervised, using EMR data.Materials and methods We classified laboratory data (arterial blood gas (ABG) measurements) and physiologic data (cardiac output (CO) measurements) derived from adult patients in the intensive care unit using decision trees and naive Bayes classifiers. Continuous features were partitioned using two supervised, and four unsupervised discretization strategies. The resulting classification accuracy was compared with that obtained with the original, continuous data.Results Supervised methods were more accurate and consistent than unsupervised, but tended to produce larger decision trees. Among the unsupervised methods, equal frequency and k-means performed well overall, while equal width was significantly less accurate.Discussion This is, we believe, the first dedicated evaluation of discretization strategies using EMR data. It is unlikely that any one discretization method applies universally to EMR data. Performance was influenced by the choice of class labels and, in the case of unsupervised methods, the number of intervals. In selecting the number of intervals there is generally a trade-off between greater accuracy and greater consistency.Conclusions In general, supervised methods yield higher accuracy, but are constrained to a single specific application. Unsupervised methods do not require class labels and can produce discretized data that can be used for multiple purposes.