Beyond Open vs. Closed: Balancing Individual Privacy and Public Accountability in Data Sharing

Beyond Open vs. Closed: Balancing Individual Privacy and Public Accountability in Data Sharing
复制标题

DOI:
10.1145/3287560.3287577
复制
发表时间:
2019-01
期刊:
Proceedings of the Conference on Fairness, Accountability, and Transparency
影响因子:
--
通讯作者:
Meg Young;Luke Rodriguez;Emilyann Keller;Feiyang Sun;Boyang Sa;Jan Whittington;Bill Howe
Meg Young;Luke Rodriguez;Emilyann Keller;Feiyang Sun;Boyang Sa;Jan Whittington;Bill Howe
中科院分区:
其他
文献类型:
--
作者:
Meg Young;Luke Rodriguez;Emilyann Keller;Feiyang Sun;Boyang Sa;Jan Whittington;Bill Howe

文献摘要

被引文献

相似文献

过于敏感而不能“开放”以供分析和重新利用的数据通常作为专有信息保持“封闭”。这种二分法破坏了使算法系统更加公平,透明和负责任的努力。政府机构需要获取专有数据来执行政策,研究人员需要评估方法,公众需要追究机构的责任;所有这些需求都必须得到满足,同时保护个人隐私和公司竞争力。在本文中,我们描述了一种由第三方公私数据信托提供的综合法律技术方法,旨在平衡这些相互竞争的利益。基本会员资格使公司和机构能够低风险地访问数据,以进行合规报告和核心方法研究,而模块化数据共享协议则支持广泛的项目和用例。除非协议中另有明确规定,否则所有数据访问最初都是通过定制的合成数据集提供给最终用户的,这些数据集提供a)强有力的隐私保障,B)消除可能暴露竞争优势的信号,以及c)消除可能强化歧视性政策的偏见,同时保持对原始数据的忠实性。我们发现,将合成数据与对原始数据的强有力的法律的保护相结合,可以在透明度、所有权、隐私和研究目标之间取得平衡。这种法律技术框架可以在各种情况下构成数据信任的基础。
Data too sensitive to be "open" for analysis and re-purposing typically remains "closed" as proprietary information. This dichotomy undermines efforts to make algorithmic systems more fair, transparent, and accountable. Access to proprietary data in particular is needed by government agencies to enforce policy, researchers to evaluate methods, and the public to hold agencies accountable; all of these needs must be met while preserving individual privacy and firm competitiveness. In this paper, we describe an integrated legal-technical approach provided by a third-party public-private data trust designed to balance these competing interests. Basic membership allows firms and agencies to enable low-risk access to data for compliance reporting and core methods research, while modular data sharing agreements support a wide array of projects and use cases. Unless specifically stated otherwise in an agreement, all data access is initially provided to end users through customized synthetic datasets that offer a) strong privacy guarantees, b) removal of signals that could expose competitive advantage, and c) removal of biases that could reinforce discriminatory policies, all while maintaining fidelity to the original data. We find that using synthetic data in conjunction with strong legal protections over raw data strikes a balance between transparency, proprietorship, privacy, and research objectives. This legal-technical framework can form the basis for data trusts in a variety of contexts.