CIF21 DIBBs: User Driven Architecture for Data Discovery
CIF21 DIBBs: User Driven Architecture for Data Discovery
批准号:
1443070
负责人:
Giridhar Manepalli
金额:
$148.49万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-04-30
中文摘要
在过去几年中,科学数据集的数量、规模和可用性都有了巨大的增长。随着科学活动变得更加数据密集和协作,跨学科研究的一个关键挑战将是发现不同的数据集,在分布式存储库和注册表中进行管理。目前,互联网上的信息发现主要是通过自动化方法进行的,其特点是网络爬行和相关算法,或者是劳动密集型的索引和分类,例如国家医学图书馆的医学文献索引。存储库中有大量的数据,只有在特定领域具有专业知识的研究人员才能知道和访问这些数据。该项目构建了一个用户驱动的数据发现架构(UDADD),该功能通过在最小的输入下建立来自不同社区的全球索引来增强科学数据集的发现。在UDADD方法中,用户操作(如数据集查询或下载)驱动全局索引的构建。通过与存储库管理员的合作,自动记录和收集这些操作。提供了两个软件插件来帮助存储库与UDADD系统进行交互。该体系结构包括基于数据集使用频率和近时性的排名技术。试验体系结构将使用DataNet Federation Consortium中的协作存储库进行演示和评估。目前,6个科学和工程团体参与了该联盟,包括海洋学、社会科学、认知科学、水文学、工程学和植物生物学等国家级项目。
英文摘要
The number, size, and availability of scientific datasets have grown enormously over the last few years. As scientific activity becomes more data intensive and collaborative, a key challenge for cross-disciplinary research will be discovery of diverse data sets, managed within distributed repositories and registries. Currently, discovery of information on the Internet is largely performed through automated approaches, characterized by web crawling and associated algorithms, or labor intensive indexing and categorization, such as the National Library of Medicine index for medical literature. There are significant amounts of data housed in repositories where only researchers with expertise in the specific field know and access the data.This project builds a user driven architecture for data discovery (UDADD), a capability that enhances discovery of scientific datasets by building a global index from diverse communities with minimal input. In the UDADD approach user actions, such as dataset queries or downloads, drive the construction of a global index. These actions are recorded and gathered automatically, through cooperation with repository managers. Two software plugins are provided to help the repositories interact with the UDADD system. The architecture includes ranking techniques based on frequency and recency of use of the datasets. The pilot architecture will be demonstrated and evaluated using cooperating repositories within the DataNet Federation Consortium. Currently, six science and engineering communities participate in the consortium, including national scale projects in oceanography, social science, cognitive science, hydrology, engineering, and plant biology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Type-Based Automation of Scientific Data Management
-
批准号:1838981
-
项目类别:Standard Grant
-
资助金额:$29.78万
-
财政年份:2018
-
负责人:Giridhar Manepalli
-
依托单位:
海外基金