课题基金 / 基金详情

III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science

III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science
III:媒介:协作研究:DataHub - 数据科学协作数据集管理平台
批准号:
1513407
负责人:
Aditya Parameswaran
金额:
$33.3万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2018-08-31

项目摘要

项目成果

Aditya Parameswaran的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The rise of the Internet, smart phones, and wireless sensors has resulted in a vast trove of data about all aspects of our lives, from our social interactions to our personal preferences to our vital signs and medical records. Increasingly, "data science" teams want to collaboratively analyze these datasets, to understand trends and to extract actionable business, scientific, or social insights. Unfortunately, while there exist tools to support data analysis, much-needed underlying infrastructure and data management capabilities are missing. To this end, "DataHub", a collaborative platform for cleaning, storing, understanding, sharing, and publishing datasets, will be developed. DataHub will be a publicly accessible platform that will host private user datasets as well as public datasets retrieved from online sources. DataHub will serve as the common substrate for data science, freeing up end users from tedious dataset book-keeping tasks, and instead supporting them in their search for useful insights. DataHub will be deployed on a large scale at MIT; partnerships with organizations and groups from a variety of sectors will be leveraged upon to show benefits for real data scientists and to ensure that the proposed techniques meet real-world big data challenges. The curriculum development part of this project will lead to the training of new data scientists, and the project will also provide opportunities for graduate and undergraduate students to participate in research and learn how to do collaborative research.Unlike most systems that focus on improving performance or on supporting even more sophisticated analyses, DataHub will instead focus on simplifying and automating many fundamental book-keeping operations that are a pre-requisite to data science. Key features of DataHub will include: (1) a flexible, source code control-like versioning system for data, that efficiently branches, merges, and differences datasets; (2) new data ingest, cleaning, and wrangling tools designed to automate data cleaning process; (3) the ability to search for "related" tables and to integrate them into the analysis process; and (4) the ability to selectively share and collaborate on data sets across users and teams. Overall, DataHub will significantly reduce the amount of effort involved on the part of data scientists for preparing, analyzing, sharing, and managing data.For more information, see the project website at: http://data-hub.org
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Optimally Leveraging Density and Locality for Exploratory Browsing and Sampling
最佳地利用密度和位置进行探索性浏览和采样
DOI: 10.1145/3209900.3209903
发表时间: 2018
期刊: SIGMOD HILDA (Human-in-the-loop Data Analytics
影响因子: --
作者: [Kim, Albert, Xu, Liqi, Siddiqui, Tarique, Huang, Silu, Madden, Samuel, Parameswaran, Aditya]
通讯作者: Parameswaran, Aditya
FW-HTF-R: Human-Machine Teaming for Effective Data Work at Scale: Upskilling Defense Lawyers Working with Police and Court Process Data
  • 批准号:
    2129008
  • 项目类别:
    Standard Grant
  • 资助金额:
    $200.0万
  • 财政年份:
    2021
  • 负责人:
    Aditya Parameswaran
  • 依托单位:
AitF: Collaborative Research: Fast, Accurate, and Practical: Adaptive Sublinear Algorithms for Scalable Visualization
  • 批准号:
    1940759
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.87万
  • 财政年份:
    2019
  • 负责人:
    Aditya Parameswaran
  • 依托单位:
CAREER: Advancing Open-Ended Crowdsourcing: The Next Frontier in Crowdsourced Data Management
  • 批准号:
    1940757
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $41.34万
  • 财政年份:
    2019
  • 负责人:
    Aditya Parameswaran
  • 依托单位:
AitF: Collaborative Research: Fast, Accurate, and Practical: Adaptive Sublinear Algorithms for Scalable Visualization
海外基金