CIF21 DIBBs: User Driven Architecture for Data Discovery
CIF21 DIBBs: User Driven Architecture for Data Discovery
批准号:
1443070
负责人:
Giridhar Manepalli
金额:
$148.49万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2018-04-30
中文摘要
在过去的几年里,科学数据集的数量、规模和可用性都有了巨大的增长。 随着科学活动变得更加数据密集和协作,跨学科研究的一个关键挑战将是发现在分布式存储库和注册表中管理的不同数据集。 目前,互联网上的信息发现主要是通过自动化方法来执行的,其特征在于网络爬行和相关算法,或劳动密集型索引和分类,例如国家医学图书馆医学文献索引。 有大量的数据存放在存储库中,只有在特定领域具有专业知识的研究人员才知道和访问数据。该项目构建了一个用户驱动的数据发现架构(UDADD),该架构通过以最少的投入构建来自不同社区的全球索引来增强科学数据集的发现能力。 在UDADD方法中,用户操作(如数据集查询或下载)驱动全局索引的构建。通过与存储库管理员的合作,这些操作被自动记录和收集。提供了两个软件插件,以帮助储存库与UDADD系统进行交互。该架构包括基于数据集使用频率和新近度的排名技术。将使用数据网联合会内的合作储存库来演示和评估试验性体系结构。 目前,六个科学和工程界参与了该联盟,包括海洋学、社会科学、认知科学、水文学、工程和植物生物学方面的国家级项目。
英文摘要
The number, size, and availability of scientific datasets have grown enormously over the last few years. As scientific activity becomes more data intensive and collaborative, a key challenge for cross-disciplinary research will be discovery of diverse data sets, managed within distributed repositories and registries. Currently, discovery of information on the Internet is largely performed through automated approaches, characterized by web crawling and associated algorithms, or labor intensive indexing and categorization, such as the National Library of Medicine index for medical literature. There are significant amounts of data housed in repositories where only researchers with expertise in the specific field know and access the data.This project builds a user driven architecture for data discovery (UDADD), a capability that enhances discovery of scientific datasets by building a global index from diverse communities with minimal input. In the UDADD approach user actions, such as dataset queries or downloads, drive the construction of a global index. These actions are recorded and gathered automatically, through cooperation with repository managers. Two software plugins are provided to help the repositories interact with the UDADD system. The architecture includes ranking techniques based on frequency and recency of use of the datasets. The pilot architecture will be demonstrated and evaluated using cooperating repositories within the DataNet Federation Consortium. Currently, six science and engineering communities participate in the consortium, including national scale projects in oceanography, social science, cognitive science, hydrology, engineering, and plant biology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Type-Based Automation of Scientific Data Management
-
批准号:1838981
-
项目类别:Standard Grant
-
资助金额:$29.78万
-
财政年份:2018
-
负责人:Giridhar Manepalli
-
依托单位:
海外基金