课题基金 / 基金详情

III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science

III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science
III:媒介:协作研究:DataHub - 数据科学协作数据集管理平台
批准号:
1513443
负责人:
Samuel Madden
金额:
$33.33万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31

项目摘要

项目成果

Samuel Madden的其他基金

相似基金

相关文献

中文摘要
翻译
互联网、智能手机和无线传感器的兴起,产生了关于我们生活方方面面的海量数据,从我们的社交互动到我们的个人偏好,再到我们的生命体征和医疗记录。越来越多的“数据科学”团队希望协作分析这些数据集,了解趋势并提取可操作的商业、科学或社会见解。不幸的是,虽然存在支持数据分析的工具,但缺少急需的底层基础设施和数据管理功能。为此,将开发“数据中心”,这是一个用于清理、存储、理解、共享和发布数据集的协作平台。DataHub将是一个公众可访问的平台,将托管私人用户数据集以及从在线来源检索的公共数据集。DataHub将作为数据科学的共同基础,将最终用户从乏味的数据集簿记任务中解放出来,转而支持他们寻找有用的见解。数据中心将在麻省理工学院大规模部署;将利用与来自不同部门的组织和团体的合作伙伴关系,为真正的数据科学家展示好处,并确保拟议的技术满足现实世界的大数据挑战。该项目的课程开发部分将培养新的数据科学家,该项目还将为研究生和本科生提供参与研究和学习如何进行协作研究的机会。与大多数专注于提高性能或支持更复杂分析的系统不同,DataHub将专注于简化和自动化许多基本的簿记操作,这些操作是数据科学的先决条件。DataHub的主要功能将包括:(1)灵活的、类似源代码控制的数据版本控制系统,可有效地对数据集进行分支、合并和差异处理;(2)旨在自动化数据清理过程的新数据获取、清理和争论工具;(3)搜索“相关”表并将其集成到分析过程中的能力;以及(4)跨用户和团队有选择地共享和协作数据集的能力。总体而言,DataHub将显著减少数据科学家准备、分析、共享和管理数据的工作量。有关更多信息,请参阅该项目网站:http://data-hub.org
英文摘要
The rise of the Internet, smart phones, and wireless sensors has resulted in a vast trove of data about all aspects of our lives, from our social interactions to our personal preferences to our vital signs and medical records. Increasingly, "data science" teams want to collaboratively analyze these datasets, to understand trends and to extract actionable business, scientific, or social insights. Unfortunately, while there exist tools to support data analysis, much-needed underlying infrastructure and data management capabilities are missing. To this end, "DataHub", a collaborative platform for cleaning, storing, understanding, sharing, and publishing datasets, will be developed. DataHub will be a publicly accessible platform that will host private user datasets as well as public datasets retrieved from online sources. DataHub will serve as the common substrate for data science, freeing up end users from tedious dataset book-keeping tasks, and instead supporting them in their search for useful insights. DataHub will be deployed on a large scale at MIT; partnerships with organizations and groups from a variety of sectors will be leveraged upon to show benefits for real data scientists and to ensure that the proposed techniques meet real-world big data challenges. The curriculum development part of this project will lead to the training of new data scientists, and the project will also provide opportunities for graduate and undergraduate students to participate in research and learn how to do collaborative research.Unlike most systems that focus on improving performance or on supporting even more sophisticated analyses, DataHub will instead focus on simplifying and automating many fundamental book-keeping operations that are a pre-requisite to data science. Key features of DataHub will include: (1) a flexible, source code control-like versioning system for data, that efficiently branches, merges, and differences datasets; (2) new data ingest, cleaning, and wrangling tools designed to automate data cleaning process; (3) the ability to search for "related" tables and to integrate them into the analysis process; and (4) the ability to selectively share and collaborate on data sets across users and teams. Overall, DataHub will significantly reduce the amount of effort involved on the part of data scientists for preparing, analyzing, sharing, and managing data.For more information, see the project website at: http://data-hub.org
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Elements: A Self-tuning Anomaly Detection Service
III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures
BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
ACM SIGMOD 2012 Student Programming Contest: A Multidimensional Indexing System
海外基金