课题基金 / 基金详情

CAREER: Web Information Extraction: Integration and Scaling

CAREER: Web Information Extraction: Integration and Scaling
职业:Web 信息提取:集成和扩展
批准号:
1351029
负责人:
Douglas Downey
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2020-08-31

项目摘要

项目成果

Douglas Downey的其他基金

相似基金

相关文献

中文摘要
翻译
本计画研究网路资讯撷取,即从网际网路中自动撷取电脑可理解的知识库。 该项目解决了WIE中的两个关键挑战。首先,学术界和工业界的许多不同团队都在追求WIE,但他们缺乏将他们的知识库组合成一个更强大的整体的方法。 该项目探讨如何在WIE系统和方法中自动整合知识。其次,WIE的一个长期目标是构建可以扩展到数十亿事实的系统,并随着时间的推移不断改进。该项目正在研究新的方法,不断优化WIE系统与有限的人为干预。该项目的目标是扩展和集成WIE系统,以满足研究社区、计算行业和公众的需求。允许不同WIE系统无缝交换知识的方法可以大大加快学术界和工业界目前正在进行的Web提取工作的进展。对于公众来说,Web提取的进步有望使改进的搜索引擎能够帮助用户完成任务并回答复杂的问题。此外,通过应用程序原型,该项目将提供面向公众的信息检索工具,承诺帮助用户检索,理解和分析网络的知识更快。该项目的研究还与一项教育计划相结合,该计划包括向代表性不足的群体进行宣传,该项目采用的技术解决方案利用自然语言的概率分布。对于集成的挑战,该项目正在开发新的应用程序编程接口(API),利用自然语言的表现力来自动集成当前和未来的WIE系统,即使系统从不同类型的语料库中提取并以不同的方式表示知识。 对于扩展挑战,该项目正在开发方法来不断优化Web上文本的新统计语言模型(SLM)。该项目从理论上研究了WIE的SLM方法,询问不同的SLM可以编码哪些类型的知识,以及需要多少文本才能获得知识。此外,该项目还引入了新的SLM功能,包括扩展到更大语料库和更多语义类的方法,以及包含搭配、定量属性、意义消歧和主动选择的人类输入的新模型。该项目网站(http://websail.eecs.northwestern.edu/wie/)提供了更多的信息和结果,包括软件、语料库和评价数据集。
英文摘要
This project studies Web Information Extraction (WIE), the task of automatically extracting computer-understandable knowledge bases (KBs) from the World Wide Web. The project addresses two key challenges in WIE. First, many different teams in academia and industry are pursuing WIE, but they lack methods for combining their KBs into a more powerful whole. This project explores how to integrate knowledge automatically across WIE systems and approaches. Secondly, a long-standing goal for WIE is to construct systems that can scale to billions of facts, by continually improving themselves over time. This project is investigating new methods that continually optimize a WIE system with limited human intervention. The project's goal of scaling and integrating WIE systems promises to address needs in the research community, the computing industry, and the public. Methods that allow different WIE systems to seamlessly exchange knowledge could dramatically hasten the progress of Web extraction efforts currently underway in academia and industry. For the public, advances in Web extraction promise to enable improved search engines that can assist users with tasks and answer complex questions. Further, through application prototypes, the project will provide public-facing information retrieval tools that promise to help users retrieve, understand, and analyze the Web's knowledge more rapidly. The project's research is also integrated with an education plan that includes outreach to underrepresented groups.The technical solutions pursued in the project utilize probability distributions over natural language. For the integration challenge, the project is developing new Application Programming Interfaces (APIs) that leverage the expressiveness of natural language to automatically integrate current and future WIE systems, even when the systems extract from different types of corpora and represent knowledge in different ways. For the scaling challenge, the project is developing ways to continually optimize new Statistical Language Models (SLMs) over text on the Web. The project investigates the SLM approach for WIE theoretically, asking what types of knowledge different SLMs can encode, and how much text is required to obtain the knowledge. Further, the project introduces new SLM capabilities, including methods for scaling to larger corpora and more semantic classes, and novel models that incorporate collocations, quantitative attributes, sense disambiguation, and actively-selected human input. The project web site (http://websail.eecs.northwestern.edu/wie/) provides additional information and access to results, including software, corpora, and evaluation data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Extracting and Representing Commonsense Knowledge Using Language Models
  • 批准号:
    2006851
  • 项目类别:
    Standard Grant
  • 资助金额:
    $47.0万
  • 财政年份:
    2020
  • 负责人:
    Douglas Downey
  • 依托单位:
RI: Medium: Collaborative Research: Learning Representations of Language for Domain Adaptation
  • 批准号:
    1065270
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2011
  • 负责人:
    Douglas Downey
  • 依托单位:
III: Small: Active Learning of Language Models for Information Extraction
  • 批准号:
    1016754
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.37万
  • 财政年份:
    2010
  • 负责人:
    Douglas Downey
  • 依托单位:
国内基金
海外基金
基于动态扩散模型与代码知识迁移的Web服务特征增强方法研究
  • 批准号:
    2026JJ80511
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    肖勇
  • 依托单位:
面向Web3D虚拟学习空间的教育智能体系统构建与应用
  • 批准号:
    2025JJ80330
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    龙艳军
  • 依托单位:
基于Web3D元宇宙的实时渲染关键技术研究和应用
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    宋三泰
  • 依托单位:
基于语义理解的多轮多约束Web服务推荐技术
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位: