The Price of Fair PCA: One Extra Dimension

The Price of Fair PCA: One Extra Dimension
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
--
影响因子:
--
通讯作者:
S. Samadi;U. Tantipongpipat;Jamie Morgenstern;Mohit Singh;S. Vempala
S. Samadi;U. Tantipongpipat;Jamie Morgenstern;Mohit Singh;S. Vempala
中科院分区:
其他
文献类型:
--
作者:
S. Samadi;U. Tantipongpipat;Jamie Morgenstern;Mohit Singh;S. Vempala

文献摘要

被引文献

相似文献

我们研究主成分分析(PCA)这种标准的降维技术是否会在不经意间对两个不同群体产生具有不同保真度的数据表示。我们在几个真实世界的数据集上表明,PCA对群体A的重建误差高于对群体B的重建误差(例如,女性与男性,或受教育程度较低与较高的个体)。即使数据集中来自A和B的样本数量相近,这种情况也可能发生。这促使我们研究对A和B保持相似保真度的降维技术。我们定义了公平主成分分析的概念,并给出了一个多项式时间算法,用于找到数据的低维表示,该表示在此度量下接近最优。最后,我们在真实世界的数据集上表明,我们的算法可用于有效地生成数据的公平低维表示。
We investigate whether the standard dimensionality reduction technique of PCA inadvertently produces data representations with different fidelity for two different populations. We show on several real-world data sets, PCA has higher reconstruction error on population A than on B (for example, women versus men or lower- versus higher-educated individuals). This can happen even when the data set has a similar number of samples from A and B. This motivates our study of dimensionality reduction techniques which maintain similar fidelity for A and B. We define the notion of Fair PCA and give a polynomial-time algorithm for finding a low dimensional representation of the data which is nearly-optimal with respect to this measure. Finally, we show on real-world data sets that our algorithm can be used to efficiently generate a fair low dimensional representation of the data.