Enabling Data Science over Open Data and Massive Data Lakes
Enabling Data Science over Open Data and Massive Data Lakes
批准号:
RGPIN-2018-06012
负责人:
Miller, Renee
金额:
$3.5万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31
中文摘要
2016年,《福布斯》评估称,数据准备工作约占数据科学家工作的80%,其中准备工作包括查找和收集数据、清理和整合数据以及管理用于数据分析的数据。他们的结论是,这也是数据科学家工作中最不愉快的部分。作为科学家,他们更愿意获得新的知识和见解。矛盾的是,如果没有原则性的数据管理和准备,这些新见解充其量也就是可疑的。由于缺乏支持原则性数据准备的工具、科学框架和数学基础,支持分析的数据准备和数据管理非常耗时且令人不快。我计划在未来五年致力于解决和帮助纠正这一赤字。作为我方法的一部分,我将重点关注开放数据,这既是因为它可用于科学研究,也是因为它对政府和社会的重要性。
今天,加拿大的大多数结构化开放数据都是CSV(逗号-分隔-值)文件,有些是JSON文件,还有一小部分是RDF文件。其中很少(如果有的话)使用遵循W3C推荐的“Data on the Web最佳实践”或其他开放数据最佳实践的模式进行发布。而且,除了没有提供信息的标记(如name10)外,这些文件很少包含属性名称。尽管在公开数据集方面做了大量努力,但开放数据出版商,如加拿大政府,除了对元数据进行简单的关键字搜索外,不提供搜索功能。在不同的数据集和发布者之间,这些元数据的质量差别很大。即使数据发布者将包括模式和其他有价值的元数据,缺乏复杂的搜索功能也会给想要使用开放数据的数据科学家带来问题。数据的异质性和不完整性为理解数据的真实结构以及如何最好地将其与其他数据相匹配并最终有效地用于数据科学带来了新的问题。
我的研究将为以下方面提出新的方法:1)发现相关数据集(例如,让数据科学家找到与她的数据相连接的所有数据集,甚至是与她的数据有意义地结合的所有数据集);2)发现开放数据上的结构的新方法和数学基础(例如,在包含未对齐或透视数据的CSV文件中);以及3)在海量公共或私人数据湖中对齐开放数据和数据的新方法。我计划将我开发的方法与基准一起开源,以帮助其他科学家开发和评估针对海量开放数据的数据准备、收集和管理解决方案。这项工作将扩大开放数据在所有级别(联邦、省和市)的社会影响范围,使这些宝贵的数据能够更容易、更有效和更有原则地用于更多目的。
英文摘要
In 2016, Forbes assessed that "data preparation accounts for about 80% of the work of data scientists" where preparation includes finding and collecting data, cleaning and integrating data, and managing data for data analysis. They concluded that this is also the least enjoyable part of a data scientist's job. As scientists, they would rather be deriving new knowledge and insights. The paradox is that without principled data management and preparation, those new insights are suspect at best. Data preparation and data management in support of analysis is so time consuming and unenjoyable because of the lack of tools, scientific frameworks, and mathematical foundations to support principled data preparation. I plan to devote the next five years to addressing and helping to correct this deficit. As part of my methodology, I will focus on open data, both because of its availability for scientific research and because of its importance to governments and society.
Today, most of Canada's structured open data is in CSV (comma--separated--value) files, with some in JSON, and a tiny amount in RDF. Little if any of these datasets are being published with a schema following the W3C recommendation "Data on the Web Best Practices" or other open data best practices. And the files rarely contain attribute names beyond uninformative tags (like name10). Despite the large effort in making datasets publicly available, open data publishers, like the Canadian government do not provide search functionality beyond simple keyword search on the metadata. This metadata varies greatly in quality across different datasets and publishers. Even if data publishers were to include schemas and other valuable metadata, the lack of sophisticated search functionality creates problems for data scientists who want to use open data. The heterogeneity and incompleteness of the data create new problems for understanding the true structure of the data and how it can best be aligned with other data and ultimately used effectively for data science.
My research will propose new methods for 1) finding relevant datasets (e.g., letting a data scientist find all datasets that join with hers or even all datasets that union meaningfully with hers) at interactive speeds over massive repositories of data; 2) new methods and mathematical foundations for the discovery of structure over open data (e.g., in CSV files that contain misaligned or pivoted data); and 3) new methods for aligning open data and data within massive public or private data lakes. I plan to make the methods I develop open source along with benchmarks for helping other scientists to develop and evaluate data preparation, collection and management solutions for massive open data. This work will extend the societal reach of Open Data at all levels (federal, provincial, and municipal), allowing this valuable data to be used more easily, in more effective and principled ways, for more purposes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:RGPIN-2018-06012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$6.99万
-
财政年份:2022
-
负责人:Miller, Renee
-
依托单位:
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:RGPIN-2018-06012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.5万
-
财政年份:2021
-
负责人:Miller, Renee
-
依托单位:
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:DGDND-2018-00017
-
项目类别:DND/NSERC Discovery Grant Supplement
-
资助金额:$2.91万
-
财政年份:2020
-
负责人:Miller, Renee
-
依托单位:
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:DGDND-2018-00017
-
项目类别:DND/NSERC Discovery Grant Supplement
-
资助金额:$2.91万
-
财政年份:2019
-
负责人:Miller, Renee
-
依托单位:
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:RGPIN-2018-06012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.5万
-
财政年份:2019
-
负责人:Miller, Renee
-
依托单位:
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:RGPIN-2018-06012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.5万
-
财政年份:2018
-
负责人:Miller, Renee
-
依托单位:
Enabling Data Science over Open Data and Massive Data Lakes
-
批准号:DGDND-2018-00017
-
项目类别:DND/NSERC Discovery Grant Supplement
-
资助金额:$2.91万
-
财政年份:2018
-
负责人:Miller, Renee
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
-
批准号:--
-
项目类别:--
-
资助金额:40万元
-
批准年份:2020
-
负责人:Vikrant Gupta
-
依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
-
批准号:61373035
-
项目类别:面上项目
-
资助金额:77.0万元
-
批准年份:2013
-
负责人:冯志勇
-
依托单位:
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data
-
批准号:31070748
-
项目类别:面上项目
-
资助金额:34.0万元
-
批准年份:2010
-
负责人:Christine Nardini
-
依托单位:
高维数据的函数型数据(functional data)分析方法
-
批准号:11001084
-
项目类别:青年科学基金项目
-
资助金额:16.0万元
-
批准年份:2010
-
负责人:周迎春
-
依托单位:
染色体复制负调控因子datA在细胞周期中的作用
-
批准号:31060015
-
项目类别:地区科学基金项目
-
资助金额:25.0万元
-
批准年份:2010
-
负责人:莫日根
-
依托单位:
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: