Automated data cleaning of paediatric anthropometric data from longitudinal electronic health records: protocol and application to a large patient cohort

Automated data cleaning of paediatric anthropometric data from longitudinal electronic health records: protocol and application to a large patient cohort
复制标题

DOI:
10.1038/s41598-020-66925-7
复制
发表时间:
2020-06-23
期刊:
影响因子:
4.6
通讯作者:
Ennis, Sarah
Ennis, Sarah
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Phan, Hang T. T.;Borca, Florina;Ennis, Sarah

文献摘要

被引文献

相似文献

医疗保健中的大数据包括从多个来源整理的具有不同程度数据质量的测量数据。这些数据需要质量控制评估,以优化临床管理的质量,并在医疗保健研究中进行强大的大规模数据分析。身高和体重数据是记录最丰富的健康统计数据之一。在电子医疗记录中,向电子记录人体测量的转变,迅速增加了测量的数量。世卫组织的指导方针指导消除基于人口的极端异常值,但缺乏工具限制了纵向人体测量的清理。我们开发和优化了一种清理儿科身高和体重数据的方案,该方案结合了使用稳健的线性回归方法进行离群值检测的方法,使用了一组手动管理的6279名患者的纵向测量数据。然后,该方案被应用于从英格兰南部一家地区教学医院的6万名儿科患者那里收集的20万份患者记录。世卫组织指南检测到生物上不可信的数据
'Big data' in healthcare encompass measurements collated from multiple sources with various degrees of data quality. These data require quality control assessment to optimise quality for clinical management and for robust large-scale data analysis in healthcare research. Height and weight data represent one of the most abundantly recorded health statistics. The shift to electronic recording of anthropometric measurements in electronic healthcare records, has rapidly inflated the number of measurements. WHO guidelines inform removal of population-based extreme outliers but an absence of tools limits cleaning of longitudinal anthropometric measurements. We developed and optimised a protocol for cleaning paediatric height and weight data that incorporates outlier detection using robust linear regression methodology using a manually curated set of 6,279 patients' longitudinal measurements. The protocol was then applied to a cohort of 200,000 patient records collected from 60,000 paediatric patients attending a regional teaching hospital in South England. WHO guidelines detected biologically implausible data in