CAREER: Web Information Extraction: Integration and Scaling
CAREER: Web Information Extraction: Integration and Scaling
批准号:
1351029
负责人:
Douglas Downey
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2020-08-31
中文摘要
本计画研究网路资讯撷取,即从网际网路中自动撷取电脑可理解的知识库。 该项目解决了WIE中的两个关键挑战。首先,学术界和工业界的许多不同团队都在追求WIE,但他们缺乏将他们的知识库组合成一个更强大的整体的方法。 该项目探讨如何在WIE系统和方法中自动整合知识。其次,WIE的一个长期目标是构建可以扩展到数十亿事实的系统,并随着时间的推移不断改进。该项目正在研究新的方法,不断优化WIE系统与有限的人为干预。该项目的目标是扩展和集成WIE系统,以满足研究社区、计算行业和公众的需求。允许不同WIE系统无缝交换知识的方法可以大大加快学术界和工业界目前正在进行的Web提取工作的进展。对于公众来说,Web提取的进步有望使改进的搜索引擎能够帮助用户完成任务并回答复杂的问题。此外,通过应用程序原型,该项目将提供面向公众的信息检索工具,承诺帮助用户检索,理解和分析网络的知识更快。该项目的研究还与一项教育计划相结合,该计划包括向代表性不足的群体进行宣传,该项目采用的技术解决方案利用自然语言的概率分布。对于集成的挑战,该项目正在开发新的应用程序编程接口(API),利用自然语言的表现力来自动集成当前和未来的WIE系统,即使系统从不同类型的语料库中提取并以不同的方式表示知识。 对于扩展挑战,该项目正在开发方法来不断优化Web上文本的新统计语言模型(SLM)。该项目从理论上研究了WIE的SLM方法,询问不同的SLM可以编码哪些类型的知识,以及需要多少文本才能获得知识。此外,该项目还引入了新的SLM功能,包括扩展到更大语料库和更多语义类的方法,以及包含搭配、定量属性、意义消歧和主动选择的人类输入的新模型。该项目网站(http://websail.eecs.northwestern.edu/wie/)提供了更多的信息和结果,包括软件、语料库和评价数据集。
英文摘要
This project studies Web Information Extraction (WIE), the task of automatically extracting computer-understandable knowledge bases (KBs) from the World Wide Web. The project addresses two key challenges in WIE. First, many different teams in academia and industry are pursuing WIE, but they lack methods for combining their KBs into a more powerful whole. This project explores how to integrate knowledge automatically across WIE systems and approaches. Secondly, a long-standing goal for WIE is to construct systems that can scale to billions of facts, by continually improving themselves over time. This project is investigating new methods that continually optimize a WIE system with limited human intervention. The project's goal of scaling and integrating WIE systems promises to address needs in the research community, the computing industry, and the public. Methods that allow different WIE systems to seamlessly exchange knowledge could dramatically hasten the progress of Web extraction efforts currently underway in academia and industry. For the public, advances in Web extraction promise to enable improved search engines that can assist users with tasks and answer complex questions. Further, through application prototypes, the project will provide public-facing information retrieval tools that promise to help users retrieve, understand, and analyze the Web's knowledge more rapidly. The project's research is also integrated with an education plan that includes outreach to underrepresented groups.The technical solutions pursued in the project utilize probability distributions over natural language. For the integration challenge, the project is developing new Application Programming Interfaces (APIs) that leverage the expressiveness of natural language to automatically integrate current and future WIE systems, even when the systems extract from different types of corpora and represent knowledge in different ways. For the scaling challenge, the project is developing ways to continually optimize new Statistical Language Models (SLMs) over text on the Web. The project investigates the SLM approach for WIE theoretically, asking what types of knowledge different SLMs can encode, and how much text is required to obtain the knowledge. Further, the project introduces new SLM capabilities, including methods for scaling to larger corpora and more semantic classes, and novel models that incorporate collocations, quantitative attributes, sense disambiguation, and actively-selected human input. The project web site (http://websail.eecs.northwestern.edu/wie/) provides additional information and access to results, including software, corpora, and evaluation data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Extracting and Representing Commonsense Knowledge Using Language Models
-
批准号:2006851
-
项目类别:Standard Grant
-
资助金额:$47.0万
-
财政年份:2020
-
负责人:Douglas Downey
-
依托单位:
RI: Medium: Collaborative Research: Learning Representations of Language for Domain Adaptation
-
批准号:1065270
-
项目类别:Continuing Grant
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:Douglas Downey
-
依托单位:
III: Small: Active Learning of Language Models for Information Extraction
-
批准号:1016754
-
项目类别:Standard Grant
-
资助金额:$18.37万
-
财政年份:2010
-
负责人:Douglas Downey
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于动态扩散模型与代码知识迁移的Web服务特征增强方法研究
-
批准号:2026JJ80511
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:肖勇
-
依托单位:
面向Web3D虚拟学习空间的教育智能体系统构建与应用
-
批准号:2025JJ80330
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:龙艳军
-
依托单位:
基于Web3D元宇宙的实时渲染关键技术研究和应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:宋三泰
-
依托单位:
基于语义理解的多轮多约束Web服务推荐技术
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
Web大数据环境下基于迁移学习的跨领域推荐研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
数据智能驱动的泛在Web应用服务质量优化方法研究
-
批准号:62102009
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:马郓
-
依托单位:
基于侧信道分析的Web站点指纹识别技术研究
-
批准号:62102084
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:顾晓丹
-
依托单位:
基于时间意图的地表覆盖Web 信息发现方法研究
-
批准号:2021JJ40721
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:侯东阳
-
依托单位:
恶劣条件下Web服务QoS预测与QoS确保的服务组合卸载方法研究
-
批准号:--
-
项目类别:面上项目
-
资助金额:58万元
-
批准年份:2021
-
负责人:夏云霓
-
依托单位:
多模态Web信息检索排序学习方法研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2021
-
负责人:耿光刚
-
依托单位: