课题基金 / 基金详情

EAGER: T2K: From Tables to Knowledge

EAGER: T2K: From Tables to Knowledge
EAGER:T2K:从表格到知识
批准号:
1250627
负责人:
Anupam Joshi
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-01 至 2016-08-31

项目摘要

项目成果

Anupam Joshi的其他基金

相似基金

相关文献

中文摘要
翻译
网络让人类变得更聪明,为人们提供了获取海量知识和事实的便捷途径。语义网具有类似地增强计算机程序和设备的能力,使它们能够访问大量的数据、事实和知识。该项目正在探索直接从电子表格、数据库关系和文档表中发现的数据中自动提取新知识,并将其表示为语义Web语言RDF中高度可互操作的链接开放数据(LOD)的可行性。提取由概率图形模型指导,该模型使用从当前LOD知识资源中挖掘的统计信息。为了展示这项研究的潜在回报,该系统被用来从从医学期刊收集的表格和从data.gov等网站收集的表格中提取知识。虽然使用W3C语义网语言RDF和OWL来表示知识,但结果也适用于其他语义数据框架,如Microdata(Search Consortium)、Freebase(Google)、Probase(Microsoft)和Open Graph(Facebook)。开源的原型软件允许其他研究人员试验自动从表格中为他们的领域产生语义丰富的数据。如果成功,这样的软件提取系统有望成为新的在线知识生态的一部分-既使用现有的LOD知识来理解表格中隐含的预期含义,又产生新的事实和知识,将成为网络的一部分。这代表着公共语义数据的广度和深度大幅增加,可以使“大数据”分析更有效。
英文摘要
The Web has made humans smarter, providing ready access to vast amounts of knowledge and facts. The Semantic Web has the capacity to similarly enhance computer programs and devices by giving them access to enormous volumes of data, facts and knowledge. This project is exploring the feasibility of automatically extracting new knowledge directly from data found in spreadsheets, database relations, and document tables and representing it as highly interoperable linked open data (LOD) in the Semantic Web language RDF. The extraction is guided by probabilistic graphical models that use statistical information mined from current LOD knowledge resources. To demonstrate the potential payoff of the research, the system is used to extract knowledge from tables collected from medical journals and tables from web sites like data.gov. While the W3C semantic web languages RDF and OWL are used to represent the knowledge, the results are applicable to other semantic data frameworks such as Microdata (Search Consortium), Freebase (Google), Probase (Microsoft) and the Open Graph (Facebook). The open sourced prototype software allows other researchers to experiment with automatically producing semantically enriched data from tables for their domains.If successful, such software extraction systems are expected to become part of a new online knowledge ecology -- both consuming existing LOD knowledge to understand the intended meaning implicit in a table and producing new facts and knowledge that will become part of Web. This represents a dramatic increase in the breadth and depth of public semantic data that can make "big data" analytics more effective.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
JST: SCC-PG: Bridging the Digital Gap and Identifying Cross-Cultural Pathways for Adoption of IoT Technologies to Support Super-Aging Societies in the U.S. and Japan
EAGER:X+CS: CS Pathways for Non CS majors
Collaborative Proposal: ITR-SemDIS: Discovering Complex Relationships in the Semantic Web
Profile Driven Architecture for Data Management in Pervasive Environments
海外基金