Fairness-Aware Instrumentation of Preprocessing~Pipelines for Machine Learning

Fairness-Aware Instrumentation of Preprocessing~Pipelines for Machine Learning
复制标题

具有公平意识的预处理仪器〜机器学习管道

DOI:
10.1145/3398730.3399194
复制
发表时间:
2020
期刊:
Workshop on Human-In-the-Loop Data Analytics (HILDA'20
影响因子:
--
通讯作者:
Schelter, Sebastian
Schelter, Sebastian
中科院分区:
--
文献类型:
--
作者:
Yang, Ke;Huang, Biao;Stoyanovich, Julia;Schelter, Sebastian

文献摘要

相似文献

在ML管道中暴露和减轻偏见是一个复杂的话题,迫切需要为数据科学家提供系统级支持。人类应该被授权调试这些管道,以控制偏差并提高数据质量和代表性。我们提出了fairDAGs,这是一个开源库,可以在ML的预处理管道中提取数据流的有向无环图(DAG)表示。该库随后使用跟踪和可视化代码对管道进行仪表化,以捕获数据分布中的变化,并在数据通过管道时识别受保护组成员身份的扭曲。我们通过对公开可用的ML管道的实验来说明fairDAG的实用性。
Surfacing and mitigating bias in ML pipelines is a complex topic, with a dire need to provide system-level support to data scientists. Humans should be empowered to debug these pipelines, in order to control for bias and to improve data quality and representativeness. We propose fairDAGs, an open-source library that extracts directed acyclic graph (DAG) representations of the data flow in preprocessing pipelines for ML. The library subsequently instruments the pipelines with tracing and visualization code to capture changes in data distributions and identify distortions with respect to protected group membership as the data travels through the pipeline. We illustrate the utility of fairDAGs, with experiments on publicly available ML pipelines.