III: Medium: Dataset Search and Ranking for Data Augmentation and Explanation
III: Medium: Dataset Search and Ranking for Data Augmentation and Explanation
批准号:
2106888
负责人:
Juliana Freire
金额:
$109.32万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-01 至 2025-08-31
中文摘要
收集和编目的有关环境、社会和民众的数据量呈爆炸式增长。此外,随着对数据透明度和开放的推动,科学家、政府和组织越来越多地在网络上提供这些数据。结合分析学和机器学习方面的进步,从理论上讲,这种日益增长的数据获取途径应该会在世界上许多最重要的科学和社会问题上取得进展。然而,由于一个核心的技术障碍,这个机会经常被错过:目前,领域专家几乎不可能从大量公开可用的信息中筛选,以发现他们特定应用所需的数据集。数据存储平台,如CKAN和Dataverse,以及数据集搜索引擎,如谷歌dataset search,旨在使共享和查找数据集变得容易。但是这些系统只支持简单的、基于关键字的查询和元数据搜索,不足以让用户正确地指定他们的信息需求。研究人员设想了一种新的数据集搜索引擎,通过支持更丰富的可查找性查询集来解锁开放数据中未开发的价值,以满足分析任务的需求,并有助于构建和改进机器学习模型。通过赋予科学家和从业者发现相关数据的能力,该项目在促进领域内和跨领域的数据重用方面具有巨大潜力。该项目将开发一些方法,其中用户的现有数据构成查询的基础,从大量数据集和属性中检索额外的相关数据。要支持这样的查询,需要克服许多技术障碍。一个主要的挑战是计算效率:这个项目将开发新的算法来快速计算和搜索数据集关系。研究人员将建立在丰富多样的工具,包括随机素描和哈希算法,并贡献新的理论分析来理解这些方法。所提供的算法将处理高度结构化的数据(例如,时空)以及通用的数字或分类数据。第二个挑战是可用性:该项目将开发新的方法来评估发现的数据关系的重要性,去除巧合或虚假的关系,以及对数据集进行排序和向最终用户展示。最后,该项目将为数据集搜索问题提供一种形式,支持基于数据集关系的广泛可查找性查询。详细介绍了高中生参与STEM相关活动的积极计划。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
There has been an explosion in the volume of data that is being collected and cataloged about the environment, society, and populace. Moreover, with the push towards transparency and open data, scientists, governments, and organizations are increasingly making these data available on the Web. Combined with advances in analytics and machine learning, such growing access to data should in theory allow for progress on many of the world’s most important scientific and societal questions. However, this opportunity is often missed due to a central technical barrier: it is currently nearly impossible for domain experts to weed through the vast amount of publicly-available information to discover datasets that are needed for their specific application. Data repository platforms, such as CKAN and Dataverse, and dataset search engines, such as Google Dataset Search, aim to make it easy to share and find datasets. But these systems only support simple, keyword-based queries and metadata search, which are insufficient for users to properly specify their information needs. The investigators envision a new kind of dataset search engine that unlocks the untapped value in open data by supporting a richer set of findability queries that cater to the needs of analytics tasks, and aid in the construction and refinement of machine learning models. By empowering scientists and practitioners with the ability to discover relevant data, the project has great potential to stimulate data reuse both within and across domains.The project will develop methods where the user’s existing data forms the basis of a query that retrieves additional, related data from a large collection of datasets and attributes. There are many technical hurdles to overcome to support such queries. One primary challenge is computational efficiency: this project will develop novel algorithms for rapidly computing and searching for dataset relationships. The investigators will build on a rich variety of tools, including randomized sketching and hashing algorithms, and contribute new theoretical analyses to understand these methods. The algorithms contributed will address both highly-structured data (e.g., spatio-temporal) as well as generic numerical or categorical data. A second challenge is usability: the project will develop novel methods for assessing the significance of discovered data relationships, for pruning out coincidental or spurious relationships, and for ranking and presenting datasets to the end-user. Finally, the project will contribute a formalism to the dataset search problem that supports a wide range of findability queries based on dataset relationships. Active plans for engagement in STEM related activities for high-school students are detailed.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Simple Analysis of Priority Sampling
优先采样的简单分析
DOI:
--
发表时间:
2024
期刊:
SIAM Symposium on Simplicity in Algorithms
影响因子:
--
作者:
[Majid Daliri, Juliana Freire, Christopher Musco, Aécio Santos, Haoxiang Zhang]
通讯作者:
Haoxiang Zhang
DOI:
10.1145/3584372.3588679
发表时间:
2023-01
期刊:
Proceedings of the 42nd ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems
影响因子:
--
作者:
[Aline Bessa;Majid Daliri;Juliana Freire;Cameron Musco;Christopher Musco;Aécio S. R. Santos;H. Zhang]
通讯作者:
Aline Bessa;Majid Daliri;Juliana Freire;Cameron Musco;Christopher Musco;Aécio S. R. Santos;H. Zhang
A Sketch-based Index for Correlated Dataset Search
用于相关数据集搜索的基于草图的索引
DOI:
10.1109/icde53745.2022.00264
发表时间:
2022
期刊:
2022 IEEE 38th International Conference on Data Engineering (ICDE
影响因子:
--
作者:
[Santos, Aecio, Bessa, Aline, Musco, Christopher, Freire, Juliana]
通讯作者:
Freire, Juliana
D-ISN/Collaborative Research: An Interdisciplinary Approach to the Discovery, Analysis, and Disruption of Wildlife Trafficking Networks
-
批准号:2146306
-
项目类别:Standard Grant
-
资助金额:$65.58万
-
财政年份:2022
-
负责人:Juliana Freire
-
依托单位:
CI-EN: Enhancing and Supporting a Community-Based Data Analysis, Visualization, and Provenance Platform
-
批准号:1405927
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2014
-
负责人:Juliana Freire
-
依托单位:
CAREER: Storing, Querying and Re-Using Provenance of Computational Tasks
-
批准号:1142013
-
项目类别:Continuing Grant
-
资助金额:$43.75万
-
财政年份:2011
-
负责人:Juliana Freire
-
依托单位:
III: EAGER: Collaborative Research: A Community Experiment Platform for Reproducibility and Generalizability
-
批准号:1139832
-
项目类别:Standard Grant
-
资助金额:$19.0万
-
财政年份:2011
-
负责人:Juliana Freire
-
依托单位:
III: EAGER: Collaborative Research: A Community Experiment Platform for Reproducibility and Generalizability
-
批准号:1050422
-
项目类别:Standard Grant
-
资助金额:$19.0万
-
财政年份:2010
-
负责人:Juliana Freire
-
依托单位:
III: Medium: Provenance Analytics: Exploring Computational Tasks and their History
-
批准号:0905385
-
项目类别:Standard Grant
-
资助金额:$95.75万
-
财政年份:2009
-
负责人:Juliana Freire
-
依托单位:
CAREER: Storing, Querying and Re-Using Provenance of Computational Tasks
-
批准号:0746500
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2008
-
负责人:Juliana Freire
-
依托单位:
III-COR: Discovering and Organizing Hidden-Web Sources
-
批准号:0713637
-
项目类别:Continuing Grant
-
资助金额:$33.62万
-
财政年份:2007
-
负责人:Juliana Freire
-
依托单位:
XML Data Management: Taking Order and Updates into Account
-
批准号:0534628
-
项目类别:Continuing Grant
-
资助金额:$27.0万
-
财政年份:2006
-
负责人:Juliana Freire
-
依托单位:
CT-T: A Laboratory Workbench for Security Research
-
批准号:0524096
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Juliana Freire
-
依托单位:
Managing Complex Visualizations
-
批准号:0513692
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Juliana Freire
-
依托单位:
A Cluster Infrastructure to Support Retrieval, Management and Visualization of Massive Amounts of Data
-
批准号:0514485
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Juliana Freire
-
依托单位:
A Cluster Infrastructure to Support Retrieval, Management and Visualization of Massive Amounts of Data
-
批准号:0323604
-
项目类别:Continuing Grant
-
资助金额:$11.0万
-
财政年份:2003
-
负责人:Juliana Freire
-
依托单位:
海外基金