课题基金 / 基金详情

BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing

BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
BD Spokes:SPOKE:NORTHEAST:协作:数据共享的许可模型和生态系统
批准号:
1636766
负责人:
Samuel Madden
金额:
$44.4万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2021-12-31

项目摘要

项目成果

Samuel Madden的其他基金

相似基金

相关文献

中文摘要
翻译
共享数据集可以为行业、研究人员和非营利组织带来巨大的互惠互利。例如,公司可以从大学研究人员探索他们的数据集和发现这一事实中获利,这有助于公司改善他们的业务。与此同时,研究人员一直在寻找现实世界的数据集,以表明他们新开发的技术在实践中有效。不幸的是,在工业界和学术界的不同利益相关者之间共享相关数据集的许多尝试都失败了,或者需要大量投资才能实现数据共享。一个主要障碍是,数据往往伴随着对如何使用数据的令人望而却步的限制(例如,要求执行法律术语或其他政策、处理数据隐私问题等)。为了今天执行这些要求,律师通常参与谈判每一份合同的条款。这种为数据共享创建个人合同的过程最终以旷日持久的谈判告终,这并不少见,因为双方都在努力应对现代安全、隐私和数据共享技术的影响和可能性。更糟糕的是,担心错过数据可能被(错误)使用的漏洞,往往会阻止许多数据共享努力甚至无法开始。为了应对这些挑战,我们新的数据共享发言人将使数据提供商能够轻松地共享数据,同时对数据的使用实施限制。这项工作有两个关键组成部分:(1)创建一个数据许可模式,以促进不同组织之间不一定开放或免费的数据共享;(2)开发一个原型数据共享软件平台,该平台执行已开发许可证的条款和限制。我们相信,这些努力将对数据共享的方式产生革命性的影响。通过将数据从个人和单个组织的孤岛转移到更广泛的社会手中,我们可以解决许多重大的社会问题。这种新的数据共享话语将使数据提供商能够轻松地共享数据,同时对数据的使用实施限制。今天已经存在许多提供访问数据集的服务和平台。然而,这些平台通常提倡完全开放获取,并没有解决在处理专有数据时出现的上述问题。因此,这项工作有三个关键组成部分:(1)创建一个数据许可模式,促进不同组织之间不一定开放或免费的数据共享;(2)开发一个原型数据共享软件平台--共享数据库,它执行已开发许可证的条款和限制;以及(3)开发和整合将伴随不同许可证共享的数据集的相关元数据,使其易于搜索和解释。为了确保开发的工具和许可证有用,该项目将成立东北数据共享小组,由许多不同的利益相关者组成,使许可模式在许多应用领域(如医疗和金融)得到广泛接受和使用。这一提议的智力价值在于设计了一种许可模式和一个数据共享平台,该平台在许多不同的领域作为模板被广泛接受和使用。虽然还有其他努力来实现数据共享(例如,知识共享),但它们侧重于数据所有者愿意在互联网上公开共享数据的情况。这种许可模式和生态系统是不同的,因为它允许数据所有者执行数据共享协议中规定的某些要求(例如,允许谁访问数据),并提供工具来确保敏感信息的数据共享的安全。我们建议调查的许可证和软件将使组织更容易向适当的组织开放其数据,同时保持确保数据受到保护、访问可撤销以及访问控制和审核日志得到维护的能力。
英文摘要
Sharing of data sets can provide tremendous mutual benefits for industry, researchers and nonprofit organizations. For example, companies can profit from the fact that university researchers explore their data sets and make discoveries, which help the company to improve their business. At the same time, researchers are always on the search for real world data sets to show that their newly developed techniques work in practice. Unfortunately, many attempts to share relevant data sets between different stakeholders in industry and academia fail or require a large investment to make data sharing possible. A major obstacle is that data often comes with prohibitive restrictions on how it can be used (e.g., requiring the enforcement of legal terms or other policies, handling data privacy issues, etc.). In order to enforce these requirements today, lawyers are usually involved in negotiation the terms of each contract. It is not atypical that this process of creating an individual contract for data sharing ends up in protracted negotiations, as both sides struggle with the implications and possibilities of modern security, privacy, and data sharing techniques. Worse, fears of missing a loophole in how the data might be (mis)used often prevents many data sharing efforts from even getting started. To address these challenges, our new data sharing spoke will enable data providers to easily share data while enforcing constraints on the use of the data. This effort has two key components:(1) Creating a licensing model for data that facilitates sharing data that is not necessarily open or free between different organizations and (2) Developing a prototype data sharing software platform, ShareDB, which enforces the terms and restrictions of the developed licenses. We believe these efforts will have a transformative impact on how data sharing takes place. By moving data out of the silos of individuals and single organizations and into the hands of broader society, we can tackle many societally significant problems.This new data sharing spoke will enable data providers to easily share data while enforcing constraints on the use of the data. Many services and platforms that provide access to data sets exist already today. However, these platforms generally promote completely open access and do not address the aforementioned issues that arise when dealing with proprietary data. Thus, the effort has three key components: (1) Creating a licensing model for data that facilitates sharing data that is not necessarily open or free between different organizations, (2) developing a prototype data sharing software platform, ShareDB, which enforces the terms and restrictions of the developed licenses, and (3) developing and integrating relevant metadata that will accompany the datasets shared under the different licenses, making them easily searchable and interpretable. To ensure that the developed tools and licenses are useful, the project will form the Northeast Data Sharing Group, comprising many different stakeholders to make the licensing model widely accepted and usable in many application domains (e.g., health and finance). The intellectual merit of this proposal is to design a licensing model and a data sharing platform that is widely accepted and usable as a template in many different domains. While there exist other efforts to enable data sharing (e.g., Creative Commons), they focus on the case where the data owner is willing to openly share the data on the Internet. This licensing model and the ecosystem is different since it allows data owners to enforce certain requirements stated in a data sharing agreement (e.g., on who is allowed to access the data) and also provides tools to make data sharing of sensitive information safe. The licenses and software we propose to investigate will make it easier for organizations to open up their data to the appropriate organizations, while maintaining the ability to ensure it is protected, that access is revocable, and that access controls and audit logs are maintained.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
ATLANTIC: Making Database Differentially Private and Faster with Accuracy Guarantee
ATLANTIC:使数据库具有差分隐私性且速度更快且保证准确性
DOI: 10.14778/3476311.3476337
发表时间: 2021
期刊: Proceedings of the International Conference on Very Large Data Bases
影响因子: --
作者: [Cao, Lei, Xiao, Dongqing, Yan, Yizhou, Madden, Samuel, Li, Guoliang]
通讯作者: Li, Guoliang
DOI: 10.1145/3329859.3329877
发表时间: 2019-03
期刊: Proceedings of the Second International Workshop on Exploiting Artificial Intelligence Techniques for Data Management
影响因子: --
作者: [R. Fernandez;S. Madden]
通讯作者: R. Fernandez;S. Madden
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者: [E. Rezig;Lei Cao;Giovanni Simonini;Maxime Schoemans;S. Madden;N. Tang;M. Ouzzani;M. Stonebraker]
通讯作者: E. Rezig;Lei Cao;Giovanni Simonini;Maxime Schoemans;S. Madden;N. Tang;M. Ouzzani;M. Stonebraker
Collaborative Research: Elements: A Self-tuning Anomaly Detection Service
III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures
III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science
ACM SIGMOD 2012 Student Programming Contest: A Multidimensional Indexing System
海外基金