Large-scale and high-resolution analysis of food purchases and health outcomes

Large-scale and high-resolution analysis of food purchases and health outcomes
复制标题

DOI:
10.1140/epjds/s13688-019-0191-y
复制
发表时间:
2019-04-30
期刊:
影响因子:
3.6
通讯作者:
Del Prete, Lucia
Del Prete, Lucia
中科院分区:
计算机科学3区
文献类型:
--
作者:
Aiello, Luca Maria;Schifanella, Rossano;Del Prete, Lucia

文献摘要

被引文献

相似文献

为了补充成本高昂且规模有限的传统饮食调查,研究人员采用数字数据来推断饮食习惯对人们健康的影响。然而,在线研究的分辨率有限:它们是在国家或区域一级进行的,不能准确地捕捉所消费食物的成分。我们研究了食物消费(来自伦敦主要杂货零售商的忠诚卡)和健康结果(来自该市所有全科医生的公开医疗处方记录)之间的关联。我们分析的规模和粒度是前所未有的:我们分析了整个伦敦一年内16亿的食品购买和11亿的医疗处方。通过研究食物消费的营养水平,我们表明,营养多样性和热量是与代谢综合征相关的三种疾病患病率的两个最强预测因素:高血压,高胆固醇和糖尿病。这种综合症是一组通常与肥胖有关的症状,在富裕国家很常见,在英国每四个成年人中就有一个患有这种综合症。我们的线性回归模型在估计伦敦近1000个人口普查地区的糖尿病患病率时达到了0.6的R2,分类器可以识别(非)健康地区,准确率高达91%。有趣的是,健康地区并不一定富裕(收入的重要性低于人们的预期),并且具有独特的特征:他们倾向于系统地减少碳水化合物和糖的摄入,使营养多样化,并避免大量摄入。更一般地说,我们的研究表明,对食品杂货购买数字记录的分析可以用作健康监测的廉价和可扩展的工具,根据这些记录,从政府到保险公司到食品公司的不同利益相关者可以实施有效的预防策略。
To complement traditional dietary surveys, which are costly and of limited scale, researchers have resorted to digital data to infer the impact of eating habits on people's health. However, online studies are limited in resolution: they are carried out at country or regional level and do not capture precisely the composition of the food consumed. We study the association between food consumption (derived from the loyalty cards of the main grocery retailer in London) and health outcomes (derived from publicly-available medical prescription records of all general practitioners in the city). The scale and granularity of our analysis is unprecedented: we analyze 1.6B food item purchases and 1.1B medical prescriptions for the entire city of London over the course of one year. By studying food consumption down to the level of nutrients, we show that nutrient diversity and amount of calories are the two strongest predictors of the prevalence of three diseases related to what is called the metabolic syndrome: hypertension, high cholesterol, and diabetes. This syndrome is a cluster of symptoms generally associated with obesity, is common across the rich world, and affects one in four adults in the UK. Our linear regression models achieve an R2 of 0.6 when estimating the prevalence of diabetes in nearly 1000 census areas in London, and a classifier can identify (un)healthy areas with up to 91% accuracy. Interestingly, healthy areas are not necessarily well-off (income matters less than what one would expect) and have distinctive features: they tend to systematically eat less carbohydrates and sugar, diversify nutrients, and avoid large quantities. More generally, our study shows that analytics of digital records of grocery purchases can be used as a cheap and scalable tool for health surveillance and, upon these records, different stakeholders from governments to insurance companies to food companies could implement effective prevention strategies.