课题基金 / 基金详情

Lodie,Web Scale Information Extraction via Linked Open Data

Lodie,Web Scale Information Extraction via Linked Open Data
Lodie,通过链接开放数据提取网络规模信息
批准号:
EP/J019488/1
负责人:
Fabio Ciravegna
金额:
$68.87万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --

项目摘要

项目成果

Fabio Ciravegna的其他基金

相似基金

相关文献

中文摘要
翻译
万维网提供了对数百亿页的访问。这些页面包含的信息大部分是非结构化的,仅供人类阅读,然而我们依靠计算机“阅读”这些页面来找到我们需要的信息。拟议中的研究旨在通过实现蒂姆·伯纳斯-李最初的设想,开发技术,从根本上改善每天进行的数十亿次搜索,即网页内容可被人类和机器阅读。这种在Web最初开发期间被忽视的愿景,现在以数据Web或链接开放数据(LOD)的形式重新出现,其中数十亿条信息被链接在一起,并可用于自动化处理。然而,网页中的信息与LOD中的信息之间缺乏相互联系。许多计划,如RDFa(由W3C支持)或Microformats(由schema.org使用并由主要搜索引擎支持),都试图通过提供用LOD链接注释网页内容的能力,使机器能够理解人类可读页面中包含的信息。当前的Web信息提取(IE)技术依赖于特定领域的训练数据或通用的提取模式,通过利用LOD,提出的研究旨在开发IE方法和技术,提供普遍的、用户驱动的、Web规模的信息提取,其中IE的目标是由用户信息需求定义的,并针对覆盖无限数量领域的数十亿可用Web文档。在这项研究中,我们的目标是开发模型和算法来创建LOD和人类可读Web之间的连续体。该方法将利用LOD提供的丰富事实和有限数量的用RDFa/微格式注释的页面来学习将未注释的网页内容连接到LOD云。这将提供互惠的好处:(i)通过明确的LOD实例和概念来搜索网页,以及(ii)通过网页内容提供的丰富信息来扩展LOD。关键的挑战是开发高效的、web规模的、半监督的、迭代的学习方法,这些方法能够使用初始的“种子”数据和注释,通过生成模型来利用:(i)局部和全局信息规律(例如表中的结构化信息,以及页面和站点范围的规律);(ii)信息的冗余(或重复);(iii) LOD中可用的任何本体论限制。当学习方法从已知的互连迭代到推断新的连接时,它们必须处理由可用的文档、领域和事实的数量和种类所产生的大量噪声。除了公布研究结果和结果外,IE开发的方法还将在schema.org相关的信息提取任务(目前由b谷歌和Bing等大型搜索引擎公司推动的任务)以及国际公开评估中进行测试。作为评估的一部分,该项目将产生至少一个公开的、网络规模的IE任务(包括语料库、链接资源等),以便与其他研究人员进行研究结果的比较。该项目旨在通过探索网络规模、用户驱动任务中的信息提取,影响自然语言处理、机器学习、信息检索、网络和语义技术等领域。该项目的成功将使创建/使用LOD的新方法成为可能,并使从Web检索信息的方式发生范式转变;从对关键词的依赖转向对这些词的概念和含义(语义)的搜索和探索。
英文摘要
The World Wide Web provides access to tens of billions of pages. These pages contain information that is largely unstructured and only intended for human readability, however we are reliant on computers "reading" these pages in order to find the information we need. The proposed research intends to develop technologies to radically improve the billions of searches which are performed every day by fulfilling the initial vision, by Tim Berners-Lee, for a Web where the webpage content is readable by both humans and machines. Such a vision, disregarded during the initial development of the Web, has now come back in the form of the Web of Data, or Linked Open Data (LOD), where billions of pieces of information are linked together and made available for automated processing. There is however a lack of interconnection between the information in the webpages and that in LOD. A number of initiatives, like RDFa (supported by W3C) or Microformats (used by schema.org and supported by major search engines) are trying to enable machines to make sense of the information contained in human readable pages by providing the ability to annotate webpage content with links into LOD.While the current state of the art in Web Information Extraction (IE) relies on domain specific training data or generic extraction patterns, by leveraging LOD the proposed research aims to develop IE methodologies and technologies providing pervasive, user-driven, Web-scale information extraction where the target of the IE is defined by the user information needs and aimed at the billions of available Web documents covering an unlimited number of domains.In this research we aim to develop models and algorithms to create a continuum between LOD and the human readable Web. The approach will utilise wealth of facts available from LOD and the limited number of pages annotated with RDFa/Microformats to learn to connect unannotated webpage content to the LOD cloud. This will provide the reciprocal advantages of: (i) enabling the search of Web pages via the unambiguous LOD instances and concepts, and (ii) the extension of the LOD with the wealth of information available from webpage content.The key challenge is the development of efficient, Web-scale, semi-supervised, iterative learning methods able to use the initial "seed" data and annotations, by generating models which exploit: (i) the local and global information regularities (e.g. structured information in tables, as well as pages and site-wide regularities); (ii) the redundancy (or repetition) of information; (iii) any ontological restrictions available in LOD. As the learning methods iterate from known interconnections to infer new connections they must cope with the massive amount of noise generated by the number and variety of documents, domains and facts available.In addition to publishing the research and its findings the IE methods developed will be tested on the task of extracting information relevant to schema.org (a task currently promoted by large search engines companies such as Google and Bing) as well as in international public evaluations. As part of such evaluations the project will generate at least one publicly available, Web-scale IE task (inclusive of corpora, linked resources, etc.) to enable comparison of research results by other researchers.The project aims to impact the fields of Natural Language Processing, Machine Learning, Information Retrieval and Web and Semantic Technologies by exploring the extraction of information in Web-scale, user-driven tasks. Success in the project will enable new ways of both creating/using the LOD and providing a paradigm shift in the way information can be retrieved from the Web; away from a reliance on keywords and towards the search and exploration of the concepts and meaning (semantics) embedded in those words.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者: [F. Ciravegna;Anna Lisa Gentile;Ziqi Zhang]
通讯作者: F. Ciravegna;Anna Lisa Gentile;Ziqi Zhang
DOI: 10.1145/2479832.2479845
发表时间: 2013-06
期刊: Proceedings of the seventh international conference on Knowledge capture
影响因子: --
作者: [Anna Lisa Gentile;Ziqi Zhang;Isabelle Augenstein;F. Ciravegna]
通讯作者: Anna Lisa Gentile;Ziqi Zhang;Isabelle Augenstein;F. Ciravegna
DOI: 10.3233/sw-150180
发表时间: 2016-01-01
期刊: SEMANTIC WEB
影响因子: 3
作者: [Augenstein, Isabelle, Maynard, Diana, Ciravegna, Fabio]
通讯作者: Ciravegna, Fabio
DOI: 10.3115/v1/w14-6203
发表时间: 2014-08
期刊:
影响因子: --
作者: [Isabelle Augenstein]
通讯作者: Isabelle Augenstein
9
    RAnDMS (Real time Analysis of Digital Media Streams)
    • 批准号:
      EP/J020583/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $26.6万
    • 财政年份:
      2012
    • 负责人:
      Fabio Ciravegna
    • 依托单位:
    国内基金
    海外基金
    基于动态扩散模型与代码知识迁移的Web服务特征增强方法研究
    • 批准号:
      2026JJ80511
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2026
    • 负责人:
      肖勇
    • 依托单位:
    面向Web3D虚拟学习空间的教育智能体系统构建与应用
    • 批准号:
      2025JJ80330
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2025
    • 负责人:
      龙艳军
    • 依托单位:
    基于Web3D元宇宙的实时渲染关键技术研究和应用
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2025
    • 负责人:
      宋三泰
    • 依托单位:
    基于语义理解的多轮多约束Web服务推荐技术
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位: