课题基金 / 基金详情

BIGDATA: F: Collaborative Research: Foundations of Responsible Data Management

BIGDATA: F: Collaborative Research: Foundations of Responsible Data Management
大数据:F:协作研究:负责任的数据管理的基础
批准号:
1741022
负责人:
Hosagrahar Jagadish
金额:
$38.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2022-08-31

项目摘要

项目成果

Hosagrahar Jagadish的其他基金

相似基金

相关文献

中文摘要
翻译
大数据技术有望改善人们的生活,加快科学发现和创新,并带来积极的社会变革。然而,如果不负责任地使用,同样的技术可能会加剧不公平,限制问责,并侵犯个人隐私:不可复制的结果可能会影响全球经济政策;搜索引擎中的算法变化可能会影响选举并煽动暴力;基于有偏见数据的模型可能会使刑事司法系统中的歧视合法化并放大;算法招聘做法可能会默默地强化多样性问题,并可能违反法律;侵犯隐私和安全的行为可能会侵蚀用户的信任,并使公司面临法律和财务后果。该项目的重点是负责任地使用大数据技术--符合伦理道德规范以及法律和政策考虑。该项目为数据管理技术确立了一个基础性的新角色,在该角色中,管理整个生命周期中数据的负责任使用成为核心系统需求。该项目更广泛的目标是帮助开创数据科学的新阶段,在这个阶段,该技术不仅考虑模型的准确性,还确保其所依赖的数据尊重相关法律、社会规范和对人类的影响。该项目定义了负责任的数据管理的属性,其中包括公平(以及相关的代表性和多样性概念)、透明度(和问责制)以及数据保护。它补充了在数据挖掘和机器学习社区中所做的工作,在数据挖掘和机器学习社区中,重点是分析数据分析生命周期最后一步的公平性、问责制和透明度,并考虑可能在数据分析的上游引入的问题:在数据集选择、清理、预处理、集成和共享期间。该项目开发概念框架和算法技术,在数据使用生命周期的所有阶段支持公平、透明和数据保护特性:从数据发现和获取开始,到清理、集成、查询和最终分析。这些贡献遵循三个目标。目标1考虑负责任的数据集发现、分析和集成。目标2考虑了负责任的查询处理,并开发了一个通用框架,用于声明性规范、检查和执行公平性、代表性和多样性。AIM 3将数据保护纳入生命周期,开发促进敏感数据共享的技术,并考虑隐私和透明度之间的权衡。该项目准备围绕负责任的数据管理建立一个多学科研究议程,作为在决策和预测系统中实现公平、问责和透明度的关键因素。有关该项目的更多信息,请访问DataResponsibly.com。
英文摘要
Big Data technology promises to improve people's lives, accelerate scientific discovery and innovation, and bring about positive societal change. Yet, if not used responsibly, this same technology can reinforce inequity, limit accountability and infringe on the privacy of individuals: irreproducible results can influence global economic policy; algorithmic changes in search engines can sway elections and incite violence; models based on biased data can legitimize and amplify discrimination in the criminal justice system; algorithmic hiring practices can silently reinforce diversity issues and potentially violate the law; privacy and security violations can erode the trust of users and expose companies to legal and financial consequences. The focus of this project is on using Big Data technology responsibly -- in accordance with ethical and moral norms, and legal and policy considerations. This project establishes a foundational new role for data management technology, in which managing the responsible use of data across the lifecycle becomes a core system requirement. The broader goal of this project is to help usher in a new phase of data science, in which the technology considers not only the accuracy of the model but also ensures that the data on which it depends respect the relevant laws, societal norms, and impacts on humans. This project defines properties of responsible data management, which include fairness (and the related concepts of representativeness and diversity), transparency (and accountability), and data protection. It complements what is done in the data mining and machine learning communities, where the focus is on analyzing fairness, accountability and transparency of the final step in the data analysis lifecycle, and considers the problems that can be introduced upstream from data analysis: during dataset selection, cleaning, pre-processing, integration, and sharing. This project develops conceptual frameworks and algorithmic techniques that support fairness, transparency and data protection properties through all stages of the data usage lifecycle: beginning with data discovery and acquisition, through cleaning, integration, querying, and ultimately analysis. The contributions are structured along three aims. Aim 1 considers responsible dataset discovery, profiling, and integration. Aim 2 considers responsible query processing and develops a general framework for declarative specification, checking and enforcement of fairness, representativeness and diversity. Aim 3 incorporates data protection into the lifecycle, develops techniques to facilitate sharing of sensitive data, and considers the tradeoffs between privacy and transparency. This project is poised to establish a multidisciplinary research agenda around responsible data management as a critical factor in enabling fairness, accountability and transparency in decision-making and prediction systems. Additional information about the project is available at DataResponsibly.com.
期刊论文(24)
专著(0)
科研奖励(0)
会议论文
Patterns Count-Based Labels for Datasets
数据集基于计数的模式标签
DOI: 10.1109/icde51399.2021.00184
发表时间: 2021
期刊: {ICDE}
影响因子: --
作者: [Moskovitch, Yuval, Jagadish, H. V.]
通讯作者: Jagadish, H. V.
DOI: --
发表时间: 2019
期刊: IEEE Data Eng. Bull.
影响因子: --
作者: [Abolfazl Asudeh;H. V. Jagadish;Julia Stoyanovich]
通讯作者: Abolfazl Asudeh;H. V. Jagadish;Julia Stoyanovich
DENOUNCER: detection of unfairness in classifiers
DENOUNCER:检测分类器中的不公平性
DOI: 10.14778/3476311.3476328
发表时间: 2021
期刊: Proceedings of the VLDB Endowment
影响因子: 2.5
作者: [Li, Jinyang, Moskovitch, Yuval, Jagadish, H. V.]
通讯作者: Jagadish, H. V.
COUNTATA: Dataset Labeling Using Pattern Counts
COUNTATA:使用模式计数的数据集标记
DOI: --
发表时间: 2020
期刊: Proceedings of the VLDB Endowment
影响因子: 2.5
作者: [Moskovitch, Yuval, Jagadish, H. V.]
通讯作者: Jagadish, H. V.
共 21 条
    Collaborative Research: III: MEDIUM: Responsible Design and Validation of Algorithmic Rankers
    CIVIC-PG Track B: Understanding Native American Tribal Residents Needs through Better Data and Query Systems
    III: Medium: Collaborative Research: Fairness in Web Database Applications
    BD Hubs: Collaborative Proposal: Midwest: Midwest Big Data Hub: Building Communities to Harness the Data Revolution
    海外基金