课题基金 / 基金详情

III: Small: Domain-Agnostic Dataset Search

III: Small: Domain-Agnostic Dataset Search
III:小型:与领域无关的数据集搜索
批准号:
1816325
负责人:
Brian Davison
金额:
$51.58万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-01 至 2022-07-31

项目摘要

项目成果

Brian Davison的其他基金

相似基金

相关文献

中文摘要
翻译
今天,网络的规模如此之大,以至于人们无法想象在没有网络搜索引擎的情况下找到如此多的信息。同样,现在可用的公共数据集的数量已经变得如此之大,以至于研究人员很难在他或她的学科范围内追踪所有这些数据,而且不可能跨学科这样做。为了帮助搜索者以一种与学科无关的方式查找数据,该项目将调查新的、有希望的全内容数据集搜索方法。这项研究将提供技术并开发一种工具的原型,该工具最终可以帮助许多类型的科学家找到他们可以用来执行探索性分析和测试假说的数据。因此,这项工作将使公共数据集发现和重用成为可能,而不管数据是由谁产生的或存储在哪里。使用这些方法的数据集搜索引擎可以帮助研究人员加快工作速度,减少重复工作,从而造福社会。它还将使其他人受益,如数据记者,因为数据有望成为新的证据来源和故事发现,一种讲故事和事实核查的新方式,使报道既有意义又值得信赖。这项工作将帮助任何数据分析师定位相关数据集。该项目将对研究生和本科生的培养产生影响。这种参与将有可能扩大代表人数不足的群体的参与,并有助于编写教材。研究人员将把这项工作的成果纳入课程,包括数据科学、网络搜索引擎、数据新闻学和语义网主题。现有的数据集搜索服务很麻烦,专注于搜索描述,而不是数据,并迎合了在自己的学科范围内寻找的搜索者。该项目的目标是开发一个原型数据集搜索引擎,其中包含用于全内容索引的新技术,使搜索者能够在网络上查找数据,而不受领域的限制。研究人员将结合信息检索、数据库和数据挖掘的原理和新方法。原型的设计和开发还将采取以用户为中心的方法,让专业人员和从业人员参与观察、访谈和实验研究,以告知和指导这一过程。这项工作的成果包括:(1)开发了从数十万真实世界公共数据集中构建搜索索引的新原则、方法和技术:研究人员将创建新的方法,用于a)全内容索引和分析,b)当现有描述符缺失时推断属性名称等附加元数据,以及c)推断可用于解决模式和数据异构性的附加描述符。(2)对搜索者在搜索和考虑使用数据集时的认知过程的理解。将建立一个社会认知模型来描述数据集搜索中的人与系统的交互,并预测系统在各种场景下的有效性。(3)开发新的界面,以支持向此类用户搜索、探索和呈现数据集。通过这一过程,研究人员将开发一套工具,从用户的角度评估数据集搜索技术和界面。研究成果将通过在会议和期刊上展示和发表、在网络上分享、演讲和将开发的软件开源来广泛传播。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Today, the size of the Web is such that one cannot imagine finding much information without a web search engine. Similarly, the number of collections of public datasets now available has become so large as to be difficult for a researcher to track all of them within his or her discipline, and impossible to do so across disciplines. To help searchers find data in a discipline-agnostic manner, this project will investigate new, promising approaches to full-content dataset search. This research will provide the technology and develop the prototype of a tool that can ultimately assist many kinds of scientists to locate data that they can use to perform exploratory analysis and test hypotheses. Thus, this work will enable public dataset discovery and reuse, regardless of who produced the data or where it is stored. A dataset search engine using these methods benefits society by helping researchers to accelerate their work and reduce duplicate efforts. It will also benefit others, such as data journalists, as data promises a new source of evidence and for story discovery, a new way for story-telling and fact-checking, to make reporting that is both meaningful and trustworthy. This work will help any data analyst locate relevant datasets. This project will impact the training of graduate students and undergraduates. This involvement will make it possible to broaden participation by underrepresented groups and the development of educational materials. The researchers will incorporate results of this work in courses, including Data Science, Web Search Engines, Data Journalism, and Semantic Web Topics. Existing dataset search services are cumbersome, focusing on searching descriptions, not data, and cater to searchers looking within their own discipline. The project's goal is to develop a prototype dataset search engine incorporating new techniques for full-content indexing to enable searchers to find data across the web, regardless of domain. The investigators will combine principles and novel methods from information retrieval, databases, and data mining. The design and development of the prototype will also take a user-centric approach, involving professionals and practitioners in observational, interview and experimental studies to inform and guide this process. The outcomes of this work include: (1) The development of new principles, methods, and technologies for the construction of search indexes from hundreds of thousands of real-world public datasets: the researchers will create novel methods for a) full-content indexing and analysis, b) inferring additional metadata such as attribute names when the existing descriptors are lacking and, c) inferring additional descriptors that can be used to resolve schema and data heterogeneity. (2) The understanding of searchers' cognitive processes as they search for and consider use of datasets. A social cognitive model will be built to describe human-system interactions in dataset searches, and to predict the effectiveness of the system in various scenarios. (3) The development of novel interfaces to support the search, exploration, and presentation of datasets to such users. Through this process, the researchers will develop a set of instruments for evaluating the dataset search technology and interface from the user's perspective. Research results will be disseminated broadly by presenting and publishing at conferences and journals, sharing on the web, giving talks, and making developed software open source.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/978-3-030-45439-5_18
发表时间: 2020-03-17
期刊: Advances in Information Retrieval
影响因子: --
作者: [Chen Z, Jia H, Heflin J, Davison BD]
通讯作者: Davison BD
MGNETS: Multi-Graph Neural Networks for Table Search
MGNETS:用于表搜索的多图神经网络
DOI: 10.1145/3459637.3482140
发表时间: 2021
期刊: Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM
影响因子: --
作者: [Chen, Zhiyu, Trabelsi, Mohamed, Heflin, Jeff, Yin, Dawei, Davison, Brian D.]
通讯作者: Davison, Brian D.
DOI: 10.1007/s10791-021-09398-0
发表时间: 2021-02
期刊: Information Retrieval Journal
影响因子: 2.5
作者: [M. Trabelsi;Zhiyu Chen;Brian D. Davison;J. Heflin]
通讯作者: M. Trabelsi;Zhiyu Chen;Brian D. Davison;J. Heflin
DOI: 10.1145/3404835.3463260
发表时间: 2021-05
期刊: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子: --
作者: [Zhiyu Chen;Shuo Zhang;B. Davison]
通讯作者: Zhiyu Chen;Shuo Zhang;B. Davison
14
    III: Small: Collaborative Research: Algorithms, systems, and theories for exploiting data dependencies in crowdsourcing
    • 批准号:
      2008155
    • 项目类别:
      Standard Grant
    • 资助金额:
      $25.0万
    • 财政年份:
      2020
    • 负责人:
      Brian Davison
    • 依托单位:
    REU Site: Intelligent and Scalable Systems
    • 批准号:
      1757787
    • 项目类别:
      Standard Grant
    • 资助金额:
      $36.0万
    • 财政年份:
      2018
    • 负责人:
      Brian Davison
    • 依托单位:
    III-COR-Medium: Efficient and Effective Search Services Over Archival Webs
    • 批准号:
      0803605
    • 项目类别:
      Standard Grant
    • 资助金额:
      $90.0万
    • 财政年份:
      2008
    • 负责人:
      Brian Davison
    • 依托单位:
    CAREER: Contextual Link Analysis
    • 批准号:
      0545875
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2006
    • 负责人:
      Brian Davison
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: