Clinical decision-making and algorithmic inequality.
Clinical decision-making and algorithmic inequality.
复制标题
临床决策和算法不平等。
DOI:
10.1136/bmjqs-2022-015874
复制
发表时间:
2023
影响因子:
5.4
通讯作者:
Challen R
中科院分区:
文献类型:
--
作者:
Challen R
Decision support algorithms based on historical data will make recommendations that are influenced by past inequality. Detailed historical health data contain patterns that identify demographic factors, such as race, 1 socioeconomic status or religion. These factors are linked to societal disadvantage and hence are indirectly correlated with unequal health outcomes. Machine learning or statistical models trained on such data will be able to identify these patterns and associate unequal outcomes with these disadvantaged groups even if the demographics are not explicitly recorded in the data. 1 2 If the indirect associations later influence a decision support algorithm, it is possible to unknowingly create further disadvantage and reinforce social inequality. 2 The reinforcement of social inequality is at highest risk when the behaviour of an algorithm is not transparent, embedded in a ‘black box’and used to influence decisions in the fields of health, education, employment or justice. 3 In both machine learning and statistical models, inequality can be the result of inadequate observational data or model design. Inadequate data might be the result of practicalities in collecting data from disadvantaged groups, 3 for example, women of childbearing age are under-represented in drug trials and 96% of UK Biobank participants are of European ancestry 4 leading to underrepresentation of disadvantaged groups within the data, which even the most sophisticated models cannot correct. It may also be due to structural societal issues, such as delayed presentation to healthcare observed in lower income groups, particularly in countries without national health services, which lead to poorer health outcomes for disadvantaged groups. During model development covariates like race, gender or socioeconomic status are a poor proxy for a range of associated cultural and societal risks such as language barriers, diet, exercise, sunlight exposure, poor housing or family history, 5 about which there are typically limited data available. Although methods for identifying and correcting for complex causal relationships exist, they are entirely dependent on the existence and availability of informative data, which, we argue, remain lacking due to historical and current structural inequalities. Statistical or machine learning models developed on routinely available data cannot differentiate between these nuances, if the data are not available, and will collapse multifactorial drivers of historical poor outcomes onto non-specific factors like ethnicity. Suppose, for example, that poor outcomes in an ethnically distinct immigrant population were due to exposure to endemic disease, or some other transient socioeconomic risks, that are not well described in the data. Models trained on such historical data will not be representative of the changing risks faced by future generations, an example of ‘temporal drift’. 6 Decision support algorithms based on such models may then propagate the kind of ‘race-based’decisions predicated on historical inequality described by Vyas et al. 7