Identifying and categorizing spurious weight data in electronic medical records

Identifying and categorizing spurious weight data in electronic medical records
复制标题

DOI:
10.1093/ajcn/nqx056
复制
发表时间:
2018-03-01
影响因子:
7.1
通讯作者:
Thielke, Stephen M.
Thielke, Stephen M.
中科院分区:
医学1区
文献类型:
--
作者:
Chen, Sunny;Banks, William A.;Thielke, Stephen M.

文献摘要

被引文献

相似文献

背景:虚假权重会损害汇总度量的有效性,例如平均值和趋势。即使是体重记录中的罕见错误也会破坏电子病历(EMR)数据的实用性。目的:我们试图估计大型EMR中虚假体重值的患病率,以确定可能的原因,并开发和测试用于识别虚假体重数据的简单算法。设计:使用从VA系统中随机选择的10,000名年龄≥ 65岁的患者的EMR数据,我们检查了不同时间间隔(从1到3000天)的体重变化百分比。我们检查了描述性结果,并开发了3种算法来对体重随时间的变化程度进行分类。在分布的基础上,我们确定了最有可能是虚假的病例。我们手动检查这些数据并对错误类型进行分类。结果:数据符合预期分布。该算法可靠地识别假权重。记录中约0.8%的所有体重似乎是虚假的,并且类似于1/5的患者图表包括>= 1个虚假体重值。最常见的错误类型涉及单个数字的误输入(例如,结论:电子病历中假权值现象普遍存在。简单的算法可以识别和删除它们,从而提高EMR数据的可靠性。
Background: Spurious weights compromise the validity of summary measures, such as averages and trends. Even rare errors in weight records can undermine the utility of electronic medical record (EMR) data.Objective: We sought to estimate the prevalence of spurious weight values in a large EMR, to ascertain the likely causes, and to develop and test straightforward algorithms for identifying spurious weight data.Design: Using EMR data from 10,000 randomly selected patients aged >= 65 y in the VA system, we examined the percentage of weight change across various time intervals, from 1 to 3000 d. We examined descriptive results and developed 3 algorithms to categorize degree of weight change over time. On the basis of distributions, we identified cases that were most likely spurious. We manually reviewed these and categorized the type of error.Results: The data followed the expected distributions. The algorithms reliably identified spurious weight. Approximately 0.8% of all weights in the record appeared to be spurious and similar to 1 in 5 patient charts included >= 1 spurious weight value. The most common type of error involved the misentry of a single digit (e.g., 148 for 178).Conclusions: Spurious weights are common in EMRs. Straightforward algorithms can identify and remove them, and thus enhance the reliability of EMR data.