课题基金 / 基金详情

CRI: Global-scale Data Sharing using Statistics and Probabilities

CRI: Global-scale Data Sharing using Statistics and Probabilities
CRI:使用统计和概率进行全球范围的数据共享
批准号:
0454425
负责人:
Dan Suciu
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-07-15 至 2009-06-30

项目摘要

项目成果

Dan Suciu的其他基金

相似基金

相关文献

中文摘要
翻译
摘要提议:CNS 0454425PI:Dan Suciu研究所:华盛顿大学计划:NSF 04-588简明计算研究基础结构标题:CRI:使用统计和概率进行全球范围的数据共享这个项目将通过探索新技术对海量数据的可扩展性来解决大规模数据集成中出现的语义异构性问题。将考虑两种这样的技术。一种是基于语料库的模式匹配,其中存储、分析和预处理大量模式集合(语料库),以增强自动模式匹配。第二种技术是基于概率的查询应答,它可以高效地在概率数据库上计算复杂的SQL查询。为了研究这些技术对大规模数据集成任务的可伸缩性,将下载Web的大量片段,并将其本地存储在服务器集群上。数据实例及其模式将自动从这些Web页面中提取。生成的模式语料库将使用各种技术进行匹配,匹配结果将以概率方式解释。生成的数据组织称为语义缓存。用户将能够在语义缓存上制定丰富的查询,例如使用像SQL这样的语言。每个查询都将根据全局数据进行评估,并给予概率解释。答案将被返回给根据其概率排名的用户。该项目如果成功,可能会影响目前无法实现大规模数据集成的各种应用程序,例如科学数据共享、电子商务和应急管理系统。
英文摘要
AbstractProposal: CNS 0454425PI: Dan SuciuInstitution: University of WashingtonProgram: NSF 04-588 CISE Computing Research InfrastructureTitle: CRI: Global-scale Data Sharing using Statistics and Probabilities This project will address the problem of semantic heterogeneity that occurs in large-scale data integration by exploring the scalability of novel techniques to very large amounts of data. Two such techniques will be considered. One is corpus-based schema matching, where a large collection (corpus) of schemas is stored, analyzed, and preprocessed in order to enhance automatic schema matching. The second technique consists of probabilistic-based query answering, which efficiently computes complex SQL queries on probabilistic databases. To study the scalability of these techniques to large-scale data integration tasks, a significant fragment of the Web will be downloaded, and stored locally, on a cluster of servers. Data instances and their schemas will be extracted automatically from these Web pages. The resulting corpus of schemas will be matched using a variety of techniques, and the matches interpreted probabilistically.The resulting data organization is called the semantic cache. Users will be able to formulate rich queries over the semantic cache, for example in a language like SQL. Each query will be evaluated on the global data, and given a probabilistic interpretation. The answers will be returned to a user ranked according to their probabilities. This project has potential, if successful, to impact a variety of applications where large scale data integration is currently impossible to achieve, such as from scientific data sharing, electronic commerce, and emergency management systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: Datalog with Aggregates: Complexity, Optimization, Evaluation
  • 批准号:
    2314527
  • 项目类别:
    Standard Grant
  • 资助金额:
    $60.0万
  • 财政年份:
    2023
  • 负责人:
    Dan Suciu
  • 依托单位:
NSF-BSF: III: Small: Data Driven Schema
  • 批准号:
    2109922
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2021
  • 负责人:
    Dan Suciu
  • 依托单位:
III: Medium: Collaborative Research: Reasoning about Optimizers for Data-Intensive Systems
  • 批准号:
    1954222
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2020
  • 负责人:
    Dan Suciu
  • 依托单位:
III:Small: Optimal Query Processing meets Information Theory: from Proofs to Algorithms
  • 批准号:
    1907997
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2019
  • 负责人:
    Dan Suciu
  • 依托单位:
国内基金
海外基金
Identification and quantification of primary phytoplankton functional types in the global oceans from hyperspectral ocean color remote sensing
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    160万元
  • 批准年份:
    2022
  • 负责人:
    李忠平
  • 依托单位:
磁层亚暴触发过程的全球(global)MHD-Hall数值模拟