课题基金 / 基金详情

Encyclopedic Lexical Representations for Natural Language Processing

Encyclopedic Lexical Representations for Natural Language Processing
自然语言处理的百科全书式词汇表示
批准号:
EP/V025961/1
负责人:
Steven Schockaert
金额:
$76.1万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

Steven Schockaert的其他基金

相似基金

相关文献

中文摘要
翻译
在过去十年中,自然语言处理(NLP)领域取得了前所未有的进步,这得益于越来越强大的神经网络模型的引入。这些模型具有在训练示例中发现模式的能力,并将这些模式转移到以前看不见的测试用例中。然而,尽管他们在许多NLP任务中表现出色,但他们“理解”语言的程度仍然非常有限。关键的潜在问题是,语言理解需要大量的世界知识,而目前的NLP系统在很大程度上缺乏这些知识。在这个项目中,我们专注于概念知识,更特别的是:(i)捕捉与给定概念相关的属性(例如狮子是危险的,船可以漂浮);(ii)描述不同概念之间的关系(例如扫帚用于清洁,蜜蜂生产蜂蜜)。我们提出的方法依赖于维基百科包含大量此类知识的事实。然而,一个关键的问题是,重要的属性和关系往往没有在文本中明确提到,特别是如果它们直接从其他信息中得出,对于人类读者来说(例如,如果X是一种能飞的动物,那么X可能有翅膀)。除了学习如何提取文本中表达的知识外,我们还必须学习如何对概念知识进行推理,核心问题是概念知识应该如何表示。当前的NLP系统严重依赖于向量表示。然后,每个概念由单个向量表示。现在已经很好地理解了如何学习这种表示,并且可以直接将其纳入神经网络架构。然而,它们也有重要的理论局限性,它们可以捕获什么知识,它们只允许肤浅和启发式的推理形式。相比之下,在符号AI中,概念知识通常使用事实和规则表示。这使得强大的推理形式成为可能,但符号表示更难学习,也更难在神经网络中使用。此外,符号表示也是有限的,因为他们不能捕捉知识的程度(如相似性和典型性),这是特别限制性的问题时,建模常识knowledge.The解决方案,我们提出依赖于一个新的混合表示框架,它结合了矢量表示的主要优点与符号方法。特别地,我们将显式地表示属性和关系,就像在符号框架中一样,但是这些属性和关系将被编码为向量。因此,每个概念将与几个属性向量相关联,而相关概念对将与一个或多个关系向量相关联。因此,我们的向量将直观地扮演与事实在符号框架中扮演的角色相同的角色,相关的神经网络模型则扮演规则的角色。该项目的主要产出将包括一个综合资源,其中概念知识以这种混合方式编码。我们希望我们的资源将在NLP中发挥重要作用,因为概念知识对语言理解的重要性及其对现有资源的高度互补性。为了证明它的有用性,我们将集中在两个具有挑战性的应用:阅读理解和主题/趋势建模。我们还将开展三个案例研究。在一个案例研究中,我们将学习公司的表示,通过使用我们的资源以语义上有意义的方式总结公司的活动。在另一个案例研究中,我们将使用我们的资源来识别与给定主题相关的新闻故事。最后,我们将使用我们的方法来学习专利中新兴趋势的语义连贯描述。
英文摘要
The field of Natural Language Processing (NLP) has made unprecedented progress over the last decade, fuelled by the introduction of increasingly powerful neural network models. These models have an impressive ability to discover patterns in training examples, and to transfer these patterns to previously unseen test cases. Despite their strong performance in many NLP tasks, however, the extent to which they "understand" language is still remarkably limited. The key underlying problem is that language understanding requires a vast amount of world knowledge, which current NLP systems are largely lacking. In this project, we focus on conceptual knowledge, and more in particular on: (i) capturing what properties are associated with a given concept (e.g. lions are dangerous, boats can float); (ii) characterising how different concepts are related (e.g. brooms are used for cleaning, bees produce honey).Our proposed approach relies on the fact that Wikipedia contains a wealth of such knowledge. A key problem, however, is that important properties and relationships are often not explicitly mentioned in text, especially if they follow straightforwardly from other information, for a human reader (e.g. if X is an animal that can fly then X probably has wings). Apart from learning to extract knowledge expressed in text, we thus also have to learn how to reason about conceptual knowledge.A central question is how conceptual knowledge should be represented. Current NLP systems heavily rely on vector representations. Each concept is then represented by a single vector. It is now well-understood how such representations can be learned, and they are straightforward to incorporate into neural network architectures. However, they also have important theoretical limitations in terms of what knowledge they can capture, and they only allow for shallow and heuristic forms of reasoning. In contrast, in symbolic AI, conceptual knowledge is typically represented using facts and rules. This enables powerful forms of reasoning, but symbolic representations are harder to learn and to use in neural networks. Moreover, symbolic representations are also limited because they cannot capture aspects of knowledge that are matters of degree (e.g. similarity and typicality), which is especially restrictive when modelling commonsense knowledge.The solution we propose relies on a novel hybrid representation framework, which combines the main advantages of vector representations with those of symbolic methods. In particular, we will explicitly represent properties and relationships, as in symbolic frameworks, but these properties and relations will be encoded as vectors. Each concept will thus be associated with several property vectors, while pairs of related concepts will be associated with one or more relation vectors. Our vectors will thus intuitively play the same role that facts play in symbolic frameworks, with associated neural network models then playing the role of rules.The main output from this project will consist in a comprehensive resource, in which conceptual knowledge is encoded in this hybrid way. We expect that our resource will play an important role in NLP, given the importance of conceptual knowledge for language understanding and its highly complementary nature to existing resources. To demonstrate its usefulness, we will focus on two challenging applications: reading comprehension and topic/trend modelling. We will also develop three case studies. In one case study, we will learn representations of companies, by using our resource to summarise the activities of companies in a semantically meaningful way. In another case study, we will use our resource to identify news stories that are relevant to a given theme. Finally, we will use our methods to learn semantically coherent descriptions of emerging trends in patents.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3539618.3591667
发表时间: 2023-05
期刊: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子: --
作者: [N. Li;Hanane Kteich;Zied Bouraoui;Steven Schockaert]
通讯作者: N. Li;Hanane Kteich;Zied Bouraoui;Steven Schockaert
DOI: 10.18653/v1/2023.emnlp-main.725
发表时间: 2023
期刊:
影响因子: --
作者: [Chatterjee U]
通讯作者: Chatterjee U
Embeddings as epistemic states: Limitations on the use of pooling operators for accumulating knowledge
作为认知状态的嵌入:使用池算子来积累知识的限制
DOI: 10.1016/j.ijar.2023.108981
发表时间: 2023
期刊: International Journal of Approximate Reasoning
影响因子: 3.9
作者: [Schockaert S]
通讯作者: Schockaert S
Solving Hard Analogy Questions with Relation Embedding Chains
使用关系嵌入链解决困难类比问题
DOI: 10.18653/v1/2023.emnlp-main.382
发表时间: 2023
期刊:
影响因子: --
作者: [Kumar N]
通讯作者: Kumar N
共 7 条
    Reasoning about Structured Story Representations
    • 批准号:
      EP/W003309/1
    • 项目类别:
      Fellowship
    • 资助金额:
      $163.22万
    • 财政年份:
      2022
    • 负责人:
      Steven Schockaert
    • 依托单位:
    Enriching, repairing and merging taxonomies by inducing qualitative spatial representations from the web
    • 批准号:
      EP/K021788/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $12.61万
    • 财政年份:
      2013
    • 负责人:
      Steven Schockaert
    • 依托单位:
    海外基金