III: Small: Domain-Agnostic Dataset Search
III: Small: Domain-Agnostic Dataset Search
批准号:
1816325
负责人:
Brian Davison
金额:
$51.58万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-01 至 2022-07-31
中文摘要
今天,网络的规模是如此之大,以至于人们无法想象没有网络搜索引擎就能找到很多信息。同样,现在可用的公共数据集的数量已经变得如此之大,以至于研究人员很难在他或她的学科中跟踪所有这些数据,而不可能在跨学科中做到这一点。为了帮助搜索者以与学科无关的方式查找数据,该项目将研究新的、有前途的全内容数据集搜索方法。这项研究将提供技术和开发一个工具的原型,最终可以帮助各种科学家找到他们可以用来进行探索性分析和测试假设的数据。因此,这项工作将使公共数据集的发现和重用成为可能,而不管数据是谁产生的,也不管数据存储在哪里。使用这些方法的数据集搜索引擎通过帮助研究人员加速他们的工作和减少重复的工作来造福社会。它也将使数据记者等其他人受益,因为数据有望成为新的证据来源和故事发现(一种讲述故事和事实核查的新方式),从而使报道既有意义又值得信赖。这项工作将帮助任何数据分析师找到相关的数据集。该项目将影响研究生和本科生的培养。这种参与将有可能扩大代表性不足的群体的参与和编写教育材料。研究人员将把这项工作的结果纳入课程,包括数据科学、网络搜索引擎、数据新闻和语义网络主题。现有的数据集搜索服务很麻烦,专注于搜索描述,而不是数据,并且迎合搜索者在他们自己的领域内寻找。该项目的目标是开发一个原型数据集搜索引擎,结合新的全内容索引技术,使搜索者能够在网络上找到数据,而不受领域的限制。研究人员将结合信息检索、数据库和数据挖掘的原理和新方法。原型的设计和开发也将采取以用户为中心的方法,让专业人员和从业人员参与观察、访谈和实验研究,为这一过程提供信息和指导。这项工作的成果包括:(1)为从数十万个真实世界的公共数据集构建搜索索引开发新的原则、方法和技术:研究人员将为a)全内容索引和分析创造新的方法;b)在缺乏现有描述符时推断额外的元数据,如属性名称;c)推断可用于解决模式和数据异构的其他描述符。(2)了解搜索者在搜索和考虑使用数据集时的认知过程。将建立一个社会认知模型来描述数据集搜索中的人-系统交互,并预测系统在各种场景下的有效性。(3)开发新的接口,以支持对这些用户的数据集的搜索、探索和呈现。通过这一过程,研究人员将开发一套从用户角度评估数据集搜索技术和界面的工具。研究成果将通过在会议和期刊上发表和发表、在网络上分享、进行演讲和使开发的软件开源等方式广泛传播。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Today, the size of the Web is such that one cannot imagine finding much information without a web search engine. Similarly, the number of collections of public datasets now available has become so large as to be difficult for a researcher to track all of them within his or her discipline, and impossible to do so across disciplines. To help searchers find data in a discipline-agnostic manner, this project will investigate new, promising approaches to full-content dataset search. This research will provide the technology and develop the prototype of a tool that can ultimately assist many kinds of scientists to locate data that they can use to perform exploratory analysis and test hypotheses. Thus, this work will enable public dataset discovery and reuse, regardless of who produced the data or where it is stored. A dataset search engine using these methods benefits society by helping researchers to accelerate their work and reduce duplicate efforts. It will also benefit others, such as data journalists, as data promises a new source of evidence and for story discovery, a new way for story-telling and fact-checking, to make reporting that is both meaningful and trustworthy. This work will help any data analyst locate relevant datasets. This project will impact the training of graduate students and undergraduates. This involvement will make it possible to broaden participation by underrepresented groups and the development of educational materials. The researchers will incorporate results of this work in courses, including Data Science, Web Search Engines, Data Journalism, and Semantic Web Topics. Existing dataset search services are cumbersome, focusing on searching descriptions, not data, and cater to searchers looking within their own discipline. The project's goal is to develop a prototype dataset search engine incorporating new techniques for full-content indexing to enable searchers to find data across the web, regardless of domain. The investigators will combine principles and novel methods from information retrieval, databases, and data mining. The design and development of the prototype will also take a user-centric approach, involving professionals and practitioners in observational, interview and experimental studies to inform and guide this process. The outcomes of this work include: (1) The development of new principles, methods, and technologies for the construction of search indexes from hundreds of thousands of real-world public datasets: the researchers will create novel methods for a) full-content indexing and analysis, b) inferring additional metadata such as attribute names when the existing descriptors are lacking and, c) inferring additional descriptors that can be used to resolve schema and data heterogeneity. (2) The understanding of searchers' cognitive processes as they search for and consider use of datasets. A social cognitive model will be built to describe human-system interactions in dataset searches, and to predict the effectiveness of the system in various scenarios. (3) The development of novel interfaces to support the search, exploration, and presentation of datasets to such users. Through this process, the researchers will develop a set of instruments for evaluating the dataset search technology and interface from the user's perspective. Research results will be disseminated broadly by presenting and publishing at conferences and journals, sharing on the web, giving talks, and making developed software open source.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1007/978-3-030-45439-5_18
发表时间:
2020-03-17
期刊:
Advances in Information Retrieval
影响因子:
--
作者:
[Chen Z, Jia H, Heflin J, Davison BD]
通讯作者:
Davison BD
MGNETS: Multi-Graph Neural Networks for Table Search
MGNETS:用于表搜索的多图神经网络
DOI:
10.1145/3459637.3482140
发表时间:
2021
期刊:
Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM
影响因子:
--
作者:
[Chen, Zhiyu, Trabelsi, Mohamed, Heflin, Jeff, Yin, Dawei, Davison, Brian D.]
通讯作者:
Davison, Brian D.
DOI:
10.1007/s10791-021-09398-0
发表时间:
2021-02
期刊:
Information Retrieval Journal
影响因子:
2.5
作者:
[M. Trabelsi;Zhiyu Chen;Brian D. Davison;J. Heflin]
通讯作者:
M. Trabelsi;Zhiyu Chen;Brian D. Davison;J. Heflin
DOI:
10.1145/3404835.3463260
发表时间:
2021-05
期刊:
Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
作者:
[Zhiyu Chen;Shuo Zhang;B. Davison]
通讯作者:
Zhiyu Chen;Shuo Zhang;B. Davison
An Architecture for Cell-Centric Indexing of Datasets
以细胞为中心的数据集索引架构
DOI:
--
发表时间:
2020
期刊:
CEUR workshop proceedings
影响因子:
--
作者:
[Qiu, Lixuan, Jia, Haiyan, Davison, Brian D., Heflin, Jeff]
通讯作者:
Heflin, Jeff
共 14 条
III: Small: Collaborative Research: Algorithms, systems, and theories for exploiting data dependencies in crowdsourcing
-
批准号:2008155
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2020
-
负责人:Brian Davison
-
依托单位:
REU Site: Intelligent and Scalable Systems
-
批准号:1757787
-
项目类别:Standard Grant
-
资助金额:$36.0万
-
财政年份:2018
-
负责人:Brian Davison
-
依托单位:
III-COR-Medium: Efficient and Effective Search Services Over Archival Webs
-
批准号:0803605
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2008
-
负责人:Brian Davison
-
依托单位:
CAREER: Contextual Link Analysis
-
批准号:0545875
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2006
-
负责人:Brian Davison
-
依托单位:
Understanding and Enhancing Queries
-
批准号:0328825
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Brian Davison
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: