Assessing machine learning for fair prediction of ADHD in school pupils using a retrospective cohort study of linked education and healthcare data.

Assessing machine learning for fair prediction of ADHD in school pupils using a retrospective cohort study of linked education and healthcare data.
复制标题

DOI:
10.1136/bmjopen-2021-058058
复制
发表时间:
2022-12-05
期刊:
影响因子:
2.9
通讯作者:
Downs, Johnny
Downs, Johnny
中科院分区:
医学3区
文献类型:
--
作者:
Ter-Minassian, Lucile;Viani, Natalia;Wickersham, Alice;Cross, Lauren;Stewart, Robert;Velupillai, Sumithra;Downs, Johnny

文献摘要

参考文献

相似文献

注意力缺陷多动障碍 (ADHD) 是一种常见的儿童期疾病,但常常未被识别和治疗。为了改善获得服务的机会,需要对 ADHD 高风险人群进行准确预测,以实现有效的资源分配。使用独特的链接健康和教育数据资源,我们研究了机器学习 (ML) 方法如何预测 ADHD 风险。回顾性人群队列研究。南伦敦(2007-2013)。 n=56 258 名拥有相关教育和健康数据的学生。我们使用曲线下面积 (AUC) 比较了四种 ML 模型和一种神经网络用于 ADHD 诊断的预测准确性。使用公平的预处理算法对种族和语言偏见进行加权。随机森林和逻辑回归预测模型在人群样本(AUC 分别为 0.86 和 0.86)和临床样本(AUC 0.72 和 0.70)中提供了最高的 ADHD 预测准确性。精确率-召回率曲线分析不太有利。通过公平的预处理算法有效减少了社会人口统计学偏差,且不损失准确性。使用链接的定期收集的教育和健康数据的机器学习方法提供了准确、低成本且可扩展的 ADHD 预测模型。这些方法可以帮助确定需求领域并为资源分配提供信息。引入“公平权重”可以减轻一些社会人口统计学偏见,否则这些偏见会低估少数群体内的多动症风险。
Attention deficit hyperactivity disorder (ADHD) is a prevalent childhood disorder, but often goes unrecognised and untreated. To improve access to services, accurate predictions of populations at high risk of ADHD are needed for effective resource allocation. Using a unique linked health and education data resource, we examined how machine learning (ML) approaches can predict risk of ADHD. Retrospective population cohort study. South London (2007–2013). n=56 258 pupils with linked education and health data. Using area under the curve (AUC), we compared the predictive accuracy of four ML models and one neural network for ADHD diagnosis. Ethnic group and language biases were weighted using a fair pre-processing algorithm. Random forest and logistic regression prediction models provided the highest predictive accuracy for ADHD in population samples (AUC 0.86 and 0.86, respectively) and clinical samples (AUC 0.72 and 0.70). Precision-recall curve analyses were less favourable. Sociodemographic biases were effectively reduced by a fair pre-processing algorithm without loss of accuracy. ML approaches using linked routinely collected education and health data offer accurate, low-cost and scalable prediction models of ADHD. These approaches could help identify areas of need and inform resource allocation. Introducing ‘fairness weighting’ attenuates some sociodemographic biases which would otherwise underestimate ADHD risk within minority groups.
DOI: 10.2147/ndt.s128752
发表时间: 2017
影响因子: 3.2
作者:
Fridman M;Banaschewski T;Sikirica V;Quintero J;Chen KS
通讯作者: Chen KS
DOI: 10.1007/s00787-015-0815-0
发表时间: 2016-09-01
影响因子: 6.4
作者:
Algorta, Guillermo Perez;Dodd, Alyson Lamont;Youngstrom, Eric A.
通讯作者: Youngstrom, Eric A.
DOI: 10.1177/1087054713520616
发表时间: 2017-05-01
影响因子: 3
作者:
Daley, Matthew F.;Newton, Douglas A.;Bussing, Regina
通讯作者: Bussing, Regina
DOI: 10.1136/bmjopen-2018-024355
发表时间: 2019-06-01
期刊: BMJ OPEN
影响因子: 2.9
作者:
Downs, Johnny M.;Ford, Tamsin;Hayes, Richard
通讯作者: Hayes, Richard
DOI: 10.1097/mlr.0000000000001147
发表时间: 2019-08-01
期刊: MEDICAL CARE
影响因子: 3
作者:
Basu, Sanjay;Narayanaswamy, Rajiv
通讯作者: Narayanaswamy, Rajiv