EquiTensors: Learning Fair Integrations of Heterogeneous Urban Data

EquiTensors: Learning Fair Integrations of Heterogeneous Urban Data
复制标题

DOI:
10.1145/3448016.3452777
复制
发表时间:
2021-06
期刊:
Proceedings of the 2021 International Conference on Management of Data
影响因子:
--
通讯作者:
An Yan;Bill Howe
An Yan;Bill Howe
中科院分区:
其他
文献类型:
--
作者:
An Yan;Bill Howe

文献摘要

被引文献

相似文献

神经方法是城市预测问题的最新技术,如交通资源需求,事故风险,人群流动性和公共安全。模型性能可以通过集成来自开放数据存储库的外生特性(例如,天气、房价、交通等),但是这些未经策划的源通常太嘈杂、不完整和有偏见而不能直接使用。我们建议从异构数据集中学习集成表示,称为EquiTensors,这些数据集可以在各种任务中重用。我们将数据集对齐到一致的时空域,然后描述一个基于卷积去噪自动编码器的无监督模型来学习共享表示。我们用自适应加权扩展了这个核心综合模型,以防止某些数据集主导信号。为了对抗歧视性偏见,我们使用对抗性学习来去除与敏感属性的相关性(例如,种族或收入)。23个输入数据集和4个真实的应用的实验表明,EquiTensors可以帮助减轻包含在有偏数据中的敏感信息的影响。与此同时,使用EquiTensors的应用程序优于忽略外生特征的模型,并与使用手动选择数据集的“oracle”模型竞争。
Neural methods are state-of-the-art for urban prediction problems such as transportation resource demand, accident risk, crowd mobility, and public safety. Model performance can be improved by integrating exogenous features from open data repositories (e.g., weather, housing prices, traffic, etc.), but these uncurated sources are often too noisy, incomplete, and biased to use directly. We propose to learn integrated representations, called EquiTensors, from heterogeneous datasets that can be reused across a variety of tasks. We align datasets to a consistent spatio-temporal domain, then describe an unsupervised model based on convolutional denoising autoencoders to learn shared representations. We extend this core integrative model with adaptive weighting to prevent certain datasets from dominating the signal. To combat discriminatory bias, we use adversarial learning to remove correlations with a sensitive attribute (e.g., race or income). Experiments with 23 input datasets and 4 real applications show that EquiTensors could help mitigate the effects of the sensitive information embodied in the biased data. Meanwhile, applications using EquiTensors outperform models that ignore exogenous features and are competitive with "oracle" models that use hand-selected datasets.