Outlier detection for questionnaire data in biobanks

Outlier detection for questionnaire data in biobanks
复制标题

生物样本库中问卷数据的异常值检测

DOI:
10.1093/ije/dyz012
复制
发表时间:
2019
影响因子:
7.7
通讯作者:
Tamiya Gen
Tamiya Gen
中科院分区:
医学1区
文献类型:
--
作者:
Sakurai Rieko;Ueki Masao;Makino Satoshi;Hozawa Atsushi;Kuriyama Shinichi;Takai-Igarashi Takako;Kinoshita Kengo;Yamamoto Masayuki;Tamiya Gen

文献摘要

相似文献

背景生物库越来越多地收集,处理和存储组学与更传统的流行病学信息,需要在数据清洗相当大的努力。一个有效的离群值检测方法,减少体力劳动是非常可取的。MethodWe开发了一种无监督的机器学习方法离群值检测,即kurPCA,使用主成分分析结合峰度,以确定离群值的存在。此外,本文还提出了一种新的回归调整方法,即基于系统缺失模式的数据回归调整(RAMP)。(东北医疗Megabank组织,日本)的结果表明,kurPCA和RAMP的组合可以有效地检测已知错误或不一致的模式。结论我们通过仿真和应用的结果证实我们的方法表现良好所提出的方法是有用的许多实际的分析方案。
BackgroundBiobanks increasingly collect, process and store omics with more conventional epidemiologic information necessitating considerable effort in data cleaning. An efficient outlier detection method that reduces manual labour is highly desirable.MethodWe develop an unsupervised machine-learning method for outlier detection, namely kurPCA, that uses principal component analysis combined with kurtosis to ascertain the existence of outliers. In addition, we propose a novel regression adjustment approach to improve detection, namely the regression adjustment for data by systematic missing patterns (RAMP).ResultApplication to epidemiological record data in a large-scale biobank (Tohoku Medical Megabank Organization, Japan) shows that a combination of kurPCA and RAMP effectively detects known errors or inconsistent patterns.ConclusionsWe confirm through the results of the simulation and the application that our methods showed good performance. The proposed methods are useful for many practical analysis scenarios.