Feasibility of Reidentifying Individuals in Large National Physical Activity Data Sets From Which Protected Health Information Has Been Removed With Use of Machine Learning

Feasibility of Reidentifying Individuals in Large National Physical Activity Data Sets From Which Protected Health Information Has Been Removed With Use of Machine Learning
复制标题

DOI:
10.1001/jamanetworkopen.2018.6040
复制
发表时间:
2018-12-01
期刊:
影响因子:
13.8
通讯作者:
Aswani, Anil
Aswani, Anil
中科院分区:
医学1区
文献类型:
--
作者:
Na, Liangyuan;Yang, Cong;Aswani, Anil

文献摘要

被引文献

相似文献

重要性尽管数据聚合和受保护的健康信息的删除,但人们担心从可穿戴设备收集的去识别的身体活动(PA)数据可能会被重新识别。收集或分发此类数据的组织表明,上述措施足以确保隐私。然而,没有研究,据我们所知,已经发表,证明了可能性或不可能性的重新识别这样的活动data.Objective评估的可行性,重新识别加速度计测量PA数据,其中有地理和受保护的健康信息删除,使用支持向量机(SVM)和随机森林方法从机器学习。2018年分析了2003-2004年和2005-2006年的国家健康和营养调查数据集。在自由生活环境中连续7天收集加速度计测量的PA数据。NHANES使用多阶段概率抽样设计来选择一个代表平民非机构化家庭的样本(成人和儿童)NHANES数据集包含客观测量的运动强度,如在所有行走1周期间佩戴的加速度计所记录的。主要结果和测量主要结果是随机森林和线性SVM算法匹配将人口统计学和20分钟聚合PA数据与个人特定记录编号进行比较,并测量每个机器学习算法的正确匹配百分比。(平均[SD]年龄,40.0 [20.6]岁)和2427例儿童(平均[SD]年龄,12.3 [3.4]岁),2003-2004年,4765名成人研究纳入了2005-2006年NHANES中的2539名儿童(平均[SD]年龄,45.2 [19.9]岁)和2539名儿童(平均[SD]年龄,12.1 [3.4]岁)。随机森林算法成功地重新识别了2003-2004年NHANES中4478名成年人(94.9%)和2120名儿童(87.4%)以及2005-2006年NHANES中4470名成年人(93.8%)和2172名儿童(85.5%)的人口统计学和20分钟聚合PA数据(P
IMPORTANCE Despite data aggregation and removal of protected health information, there is concern that deidentified physical activity (PA) data collected from wearable devices can be reidentified. Organizations collecting or distributing such data suggest that the aforementioned measures are sufficient to ensure privacy. However, no studies, to our knowledge, have been published that demonstrate the possibility or impossibility of reidentifying such activity data.OBJECTIVE To evaluate the feasibility of reidentifying accelerometer-measured PA data, which have had geographic and protected health information removed, using support vector machines (SVMs) and random forest methods from machine learning.DESIGN, SETTING, AND PARTICIPANTS In this cross-sectional study, the National Health and Nutrition Examination Survey (NHANES) 2003-2004 and 2005-2006 data sets were analyzed in 2018. The accelerometer-measured PA data were collected in a free-living setting for 7 continuous days. NHANES uses a multistage probability sampling design to select a sample that is representative of the civilian noninstitutionalized household (both adult and children) population of the United States.EXPOSURES The NHANES data sets contain objectively measured movement intensity as recorded by accelerometers worn during all walking for 1 week.MAIN OUTCOMES AND MEASURES The primary outcome was the ability of the random forest and linear SVM algorithms to match demographic and 20-minute aggregated PA data to individual-specific record numbers, and the percentage of correct matches by each machine learning algorithm was the measure.RESULTS A total of 4720 adults (mean [SD] age, 40.0 [20.6] years) and 2427 children (mean [SD] age, 12.3 [3.4] years) in NHANES 2003-2004 and 4765 adults (mean [SD] age, 45.2 [19.9] years) and 2539 children (mean [SD] age, 12.1 [3.4] years) in NHANES 2005-2006 were included in the study. The random forest algorithm successfully reidentified the demographic and 20-minute aggregated PA data of 4478 adults (94.9%) and 2120 children (87.4%) in NHANES 2003-2004 and 4470 adults (93.8%) and 2172 children (85.5%) in NHANES 2005-2006 (P