An automated data cleaning method for Electronic Health Records by incorporating clinical knowledge.

An automated data cleaning method for Electronic Health Records by incorporating clinical knowledge.
复制标题

DOI:
10.1186/s12911-021-01630-7
复制
发表时间:
2021-09-17
影响因子:
3.5
通讯作者:
De Moor B
De Moor B
中科院分区:
医学3区
文献类型:
--
作者:
Shi X;Prins C;Van Pottelbergh G;Mamouris P;Vaes B;De Moor B

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)数据在临床研究中的使用正在令人难以置信地增加,但数据资源的缺乏提出了数据清洗的挑战。如果数据清理可以自动完成,则可以节省时间。此外,用于其他领域数据的自动数据清理工具通常统一处理所有变量,这意味着它们不能很好地用于临床数据,因为需要考虑变量特定的信息。本文提出了一种考虑临床知识的电子病历数据自动清洗方法。 我们使用了1994-2015年期间从比利时弗兰德斯的初级保健中收集的EHR数据。我们构建了一个临床知识数据库来存储数据清理所需的所有变量特定信息。我们应用模糊搜寻来自动侦测并取代拼错的单位,并依照特定变数的转换公式来执行单位转换。然后根据临床知识对数值进行校正并检测离群值。总共清理了52个临床变量,并比较了清理过程前后的缺失值百分比(完整性)和正常范围内的值百分比(正确性)。在数据清理之前,所有变量均已100%完成。42个变量的缺失值百分比下降小于1%,9个变量下降1- 10%。只有1个变量的完整性出现大幅下降(13.36%)。所有变量在清洁后的正常范围内的值均超过50%,其中43个变量的百分比高于70%。我们提出了一个通用的方法,临床变量,实现了高度的自动化,是能够处理大规模的数据。这种方法大大提高了数据清理的效率,为非技术人员消除了技术障碍。
The use of Electronic Health Records (EHR) data in clinical research is incredibly increasing, but the abundancy of data resources raises the challenge of data cleaning. It can save time if the data cleaning can be done automatically. In addition, the automated data cleaning tools for data in other domains often process all variables uniformly, meaning that they cannot serve well for clinical data, as there is variable-specific information that needs to be considered. This paper proposes an automated data cleaning method for EHR data with clinical knowledge taken into consideration. We used EHR data collected from primary care in Flanders, Belgium during 1994–2015. We constructed a Clinical Knowledge Database to store all the variable-specific information that is necessary for data cleaning. We applied Fuzzy search to automatically detect and replace the wrongly spelled units, and performed the unit conversion following the variable-specific conversion formula. Then the numeric values were corrected and outliers were detected considering the clinical knowledge. In total, 52 clinical variables were cleaned, and the percentage of missing values (completeness) and percentage of values within the normal range (correctness) before and after the cleaning process were compared. All variables were 100% complete before data cleaning. 42 variables had a drop of less than 1% in the percentage of missing values and 9 variables declined by 1–10%. Only 1 variable experienced large decline in completeness (13.36%). All variables had more than 50% values within the normal range after cleaning, of which 43 variables had a percentage higher than 70%. We propose a general method for clinical variables, which achieves high automation and is capable to deal with large-scale data. This method largely improved the efficiency to clean the data and removed the technical barriers for non-technical people.
DOI: 10.1186/s12911-019-0740-0
发表时间: 2019-02-12
影响因子: 3.5
作者:
Terry, Amanda L.;Stewart, Moira;Thind, Amardeep
通讯作者: Thind, Amardeep
DOI: 10.13063/2327-9214.1201
发表时间: 2016
期刊: EGEMS (Washington, DC)
影响因子: --
作者:
Dziadkowiec O;Callahan T;Ozkaynak M;Reeder B;Welton J
通讯作者: Welton J
DOI: 10.2174/1874431101812010019
发表时间: 2018-01-01
期刊: The open medical informatics journal
影响因子: --
作者:
Mashoufi, Mehrnaz;Ayatollahi, Haleh;Khorasani-Zavareh, Davoud
通讯作者: Khorasani-Zavareh, Davoud
DOI: 10.1038/s41598-020-66925-7
发表时间: 2020-06-23
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Phan, Hang T. T.;Borca, Florina;Ennis, Sarah
通讯作者: Ennis, Sarah
DOI: 10.1097/mlr.0b013e318257dd67
发表时间: 2012-07
期刊: Medical care
影响因子: 3
作者:
Kahn MG;Raebel MA;Glanz JM;Riedlinger K;Steiner JF
通讯作者: Steiner JF