Encyclopedic Lexical Representations for Natural Language Processing
Encyclopedic Lexical Representations for Natural Language Processing
批准号:
EP/V025961/1
负责人:
Steven Schockaert
金额:
$76.1万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The field of Natural Language Processing (NLP) has made unprecedented progress over the last decade, fuelled by the introduction of increasingly powerful neural network models. These models have an impressive ability to discover patterns in training examples, and to transfer these patterns to previously unseen test cases. Despite their strong performance in many NLP tasks, however, the extent to which they "understand" language is still remarkably limited. The key underlying problem is that language understanding requires a vast amount of world knowledge, which current NLP systems are largely lacking. In this project, we focus on conceptual knowledge, and more in particular on: (i) capturing what properties are associated with a given concept (e.g. lions are dangerous, boats can float); (ii) characterising how different concepts are related (e.g. brooms are used for cleaning, bees produce honey).Our proposed approach relies on the fact that Wikipedia contains a wealth of such knowledge. A key problem, however, is that important properties and relationships are often not explicitly mentioned in text, especially if they follow straightforwardly from other information, for a human reader (e.g. if X is an animal that can fly then X probably has wings). Apart from learning to extract knowledge expressed in text, we thus also have to learn how to reason about conceptual knowledge.A central question is how conceptual knowledge should be represented. Current NLP systems heavily rely on vector representations. Each concept is then represented by a single vector. It is now well-understood how such representations can be learned, and they are straightforward to incorporate into neural network architectures. However, they also have important theoretical limitations in terms of what knowledge they can capture, and they only allow for shallow and heuristic forms of reasoning. In contrast, in symbolic AI, conceptual knowledge is typically represented using facts and rules. This enables powerful forms of reasoning, but symbolic representations are harder to learn and to use in neural networks. Moreover, symbolic representations are also limited because they cannot capture aspects of knowledge that are matters of degree (e.g. similarity and typicality), which is especially restrictive when modelling commonsense knowledge.The solution we propose relies on a novel hybrid representation framework, which combines the main advantages of vector representations with those of symbolic methods. In particular, we will explicitly represent properties and relationships, as in symbolic frameworks, but these properties and relations will be encoded as vectors. Each concept will thus be associated with several property vectors, while pairs of related concepts will be associated with one or more relation vectors. Our vectors will thus intuitively play the same role that facts play in symbolic frameworks, with associated neural network models then playing the role of rules.The main output from this project will consist in a comprehensive resource, in which conceptual knowledge is encoded in this hybrid way. We expect that our resource will play an important role in NLP, given the importance of conceptual knowledge for language understanding and its highly complementary nature to existing resources. To demonstrate its usefulness, we will focus on two challenging applications: reading comprehension and topic/trend modelling. We will also develop three case studies. In one case study, we will learn representations of companies, by using our resource to summarise the activities of companies in a semantically meaningful way. In another case study, we will use our resource to identify news stories that are relevant to a given theme. Finally, we will use our methods to learn semantically coherent descriptions of emerging trends in patents.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3539618.3591667
发表时间:
2023-05
期刊:
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
作者:
[N. Li;Hanane Kteich;Zied Bouraoui;Steven Schockaert]
通讯作者:
N. Li;Hanane Kteich;Zied Bouraoui;Steven Schockaert
DOI:
10.18653/v1/2023.emnlp-main.725
发表时间:
2023
期刊:
影响因子:
--
作者:
[Chatterjee U]
通讯作者:
Chatterjee U
Embeddings as epistemic states: Limitations on the use of pooling operators for accumulating knowledge
作为认知状态的嵌入:使用池算子来积累知识的限制
DOI:
10.1016/j.ijar.2023.108981
发表时间:
2023
期刊:
International Journal of Approximate Reasoning
影响因子:
3.9
作者:
[Schockaert S]
通讯作者:
Schockaert S
Solving Hard Analogy Questions with Relation Embedding Chains
使用关系嵌入链解决困难类比问题
DOI:
10.18653/v1/2023.emnlp-main.382
发表时间:
2023
期刊:
影响因子:
--
作者:
[Kumar N]
通讯作者:
Kumar N
Ultra-Fine Entity Typing with Prior Knowledge about Labels: A Simple Clustering Based Strategy
具有标签先验知识的超精细实体类型:基于简单聚类的策略
DOI:
10.18653/v1/2023.findings-emnlp.786
发表时间:
2023
期刊:
影响因子:
--
作者:
[Li N]
通讯作者:
Li N
共 7 条
Reasoning about Structured Story Representations
-
批准号:EP/W003309/1
-
项目类别:Fellowship
-
资助金额:$163.22万
-
财政年份:2022
-
负责人:Steven Schockaert
-
依托单位:
Enriching, repairing and merging taxonomies by inducing qualitative spatial representations from the web
-
批准号:EP/K021788/1
-
项目类别:Research Grant
-
资助金额:$12.61万
-
财政年份:2013
-
负责人:Steven Schockaert
-
依托单位:
海外基金