Clinical decision-making and algorithmic inequality.

Clinical decision-making and algorithmic inequality.
复制标题

临床决策和算法不平等。

DOI:
10.1136/bmjqs-2022-015874
复制
发表时间:
2023
影响因子:
5.4
通讯作者:
Challen R
Challen R
中科院分区:
医学1区
文献类型:
--
作者:
Challen R

文献摘要

相似文献

基于历史数据的决策支持算法将做出受过去不平等影响的建议。详细的历史健康数据包含识别人口统计因素的模式,例如种族、1社会经济地位或宗教。这些因素与社会不利因素有关,因此与不平等的健康结果间接相关。在这些数据上训练的机器学习或统计模型将能够识别这些模式,并将不平等的结果与这些弱势群体联系起来,即使数据中没有明确记录人口统计数据。如果间接关联后来影响了决策支持算法,则可能在不知不觉中造成进一步的劣势并加剧社会不平等。2.当算法的行为不透明、嵌入“黑匣子”并用于影响卫生、教育、就业或司法领域的决策时,社会不平等加剧的风险最高。在机器学习和统计模型中,不平等可能是观测数据或模型设计不足的结果。数据不足可能是从弱势群体收集数据的实际性造成的,3例如,育龄妇女在药物试验中的代表性不足,96%的英国生物库参与者是欧洲血统4,导致数据中弱势群体的代表性不足,即使是最复杂的模型也无法纠正。这也可能是由于结构性的社会问题,例如在低收入群体中观察到的医疗保健延迟,特别是在没有国家卫生服务的国家,这导致弱势群体的健康结果较差。在模型开发过程中,种族、性别或社会经济地位等协变量对于一系列相关的文化和社会风险(如语言障碍、饮食、锻炼、阳光照射、不良住房或家族史)来说是一个很差的替代指标,5关于这些风险,通常可用的数据有限。虽然存在识别和纠正复杂因果关系的方法,但它们完全依赖于信息数据的存在和可用性,我们认为,由于历史和当前的结构性不平等,这些数据仍然缺乏。如果数据不可用,基于常规可用数据开发的统计或机器学习模型无法区分这些细微差别,并且会将历史不良结果的多因素驱动因素分解为种族等非特定因素。例如,假设一个种族不同的移民群体的不良结果是由于暴露于地方病或其他一些暂时的社会经济风险,这些风险在数据中没有得到很好的描述。在这些历史数据上训练的模型将不能代表未来几代人面临的不断变化的风险,这是“时间漂移”的一个例子。6基于这种模型的决策支持算法可能会传播Vyas等人描述的基于历史不平等的“基于种族”的决策。
Decision support algorithms based on historical data will make recommendations that are influenced by past inequality. Detailed historical health data contain patterns that identify demographic factors, such as race, 1 socioeconomic status or religion. These factors are linked to societal disadvantage and hence are indirectly correlated with unequal health outcomes. Machine learning or statistical models trained on such data will be able to identify these patterns and associate unequal outcomes with these disadvantaged groups even if the demographics are not explicitly recorded in the data. 1 2 If the indirect associations later influence a decision support algorithm, it is possible to unknowingly create further disadvantage and reinforce social inequality. 2 The reinforcement of social inequality is at highest risk when the behaviour of an algorithm is not transparent, embedded in a ‘black box’and used to influence decisions in the fields of health, education, employment or justice. 3 In both machine learning and statistical models, inequality can be the result of inadequate observational data or model design. Inadequate data might be the result of practicalities in collecting data from disadvantaged groups, 3 for example, women of childbearing age are under-represented in drug trials and 96% of UK Biobank participants are of European ancestry 4 leading to underrepresentation of disadvantaged groups within the data, which even the most sophisticated models cannot correct. It may also be due to structural societal issues, such as delayed presentation to healthcare observed in lower income groups, particularly in countries without national health services, which lead to poorer health outcomes for disadvantaged groups. During model development covariates like race, gender or socioeconomic status are a poor proxy for a range of associated cultural and societal risks such as language barriers, diet, exercise, sunlight exposure, poor housing or family history, 5 about which there are typically limited data available. Although methods for identifying and correcting for complex causal relationships exist, they are entirely dependent on the existence and availability of informative data, which, we argue, remain lacking due to historical and current structural inequalities. Statistical or machine learning models developed on routinely available data cannot differentiate between these nuances, if the data are not available, and will collapse multifactorial drivers of historical poor outcomes onto non-specific factors like ethnicity. Suppose, for example, that poor outcomes in an ethnically distinct immigrant population were due to exposure to endemic disease, or some other transient socioeconomic risks, that are not well described in the data. Models trained on such historical data will not be representative of the changing risks faced by future generations, an example of ‘temporal drift’. 6 Decision support algorithms based on such models may then propagate the kind of ‘race-based’decisions predicated on historical inequality described by Vyas et al. 7