ConCur: Knowledge Base Construction and Curation
ConCur: Knowledge Base Construction and Curation
批准号:
EP/V050869/1
负责人:
Ian Horrocks
金额:
$144.12万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
知识图是图式结构的知识资源,通常表示为三元组,如(“UK”,“hasCapital”,“London”)和(“London”,“instanceOf”,“City”)。除了这些基本的“事实”之外,知识图通常还包括关于领域的结构化知识,通常基于实体类型(AKA类或概念)的层次结构;例如,(“City”、“subClassOf”、“HumanSetation”)。大部分或全部由结构化知识组成的知识图通常被称为本体,一些知识图是通用的,如维基数据和谷歌知识图,而另一些知识图是为特定领域开发的,如医学。它们的重要性正在迅速提高,并在许多应用中发挥着关键作用。例如,谷歌使用其知识图谱进行搜索、问答和谷歌助手,而亚马逊和苹果也分别使用知识图谱为其个人助理Alexa和Siri提供动力。知识图谱被广泛用于健康和福祉领域,例如,用于组织和交换信息,以及为临床人工智能(AI)提供动力。一个例子是FoodOn,这是一个本体,表示食品知识,如细粒食品分类、营养和过敏原,以及相关活动,如农业。然而,知识图的构建和维护非常具有挑战性,可能需要大量的人力。尽管知识创造的成本很高,但知识图谱往往仍然是有偏见的、不完整的或过于粗粒度的。以健康和生活方式的本体论Helis为例。它的食物知识非常简单,通常用一个实体代表许多不同的变体(例如,香蕉的所有种类和衍生物),与专门的生物医学本体论相比,它对健康的知识是高度不完整的。此外,在知识图谱中很难避免错误的事实和分类;例如,FoodOn将豆奶归类为一种牛奶,而不是一种豆制品。这种错误可能是从信息源继承的,也可能是由施工程序引起的。这些问题严重影响了知识图的有用性和使用知识图的系统的可靠性;例如,如果将知识图用于食物过敏原警报系统,豆奶的分类可能会有危险。因此,迫切需要有效的知识图构建和管理,并将在充分利用知识图的全部价值方面发挥关键作用。由于现在有许多可用的知识资源,一种可能的办法是使用多个来源来解决覆盖面和质量问题,例如通过整合和交叉检查。例如,将Helis与FoodOn相结合将把食品(包括香蕉)的细粒度分类与生活方式知识结合起来。此外,与Helis交叉检查FoodOn将发现豆奶的问题,在Helis,豆奶被正确归类为豆制品。知识资源的自动化整合具有挑战性,但语义和基于学习的技术相结合似乎是一种非常有前途的方法,我们已经在这个方向上取得了一些令人鼓舞的初步结果。因此,拟议的研究将研究一系列语义和机器学习技术,以及如何将它们结合起来支持知识图的构建和管理。除了在知识图谱的构建和管理中的应用,本研究还将有助于发展新的神经符号理论、范式和方法,如用于表达知识的学习表征的深度语义嵌入,以及用于解决样本不足问题的知识引导学习。这些技术有望给许多人工智能和大数据技术带来革命性的变化。
英文摘要
Knowledge graphs are graph-structured knowledge resources which are often expressed as triples such as ("UK", "hasCapital", "London") and ("London", "instanceOf", "City"). As well as such basic "facts", knowledge graphs often include structural knowledge about the domain, typically based on a hierarchy of entity types (AKA classes or concepts); e.g., ("City", "subClassOf", "HumanSettlement"). A knowledge graph that consist largely or wholly of structural knowledge is often called an ontology.Some knowledge graphs are general purpose, such as Wikidata and the Google knowledge graph, while others are developed for specific domains such as medicine. They are rapidly gaining in importance and are playing a key role in many applications. For example, Google uses its knowledge graph for search, question answering and Google Assistant, while Amazon and Apple also use knowledge graphs to power their personal assistants Alexa and Siri, respectively. Knowledge graphs are widely used in the domain of health and wellbeing, e.g., for organising and exchanging information and to power clinical artificial intelligence (AI). One example is FoodOn, an ontology representing food knowledge such as fine-grained food product categorization, nutrition and allergens, as well as related activities such as agriculture.Knowledge graph construction and maintenance is, however, very challenging, and may require a considerable amount of human effort. Notwithstanding the high cost of knowledge creation, knowledge graphs are often still biased, incomplete or too coarse-grained. Take HeLis, an ontology for health and lifestyle, as an example. Its food knowledge is quite simple and often represents many different variants with a single entity (e.g., "Banana" for all kinds and derivatives of bananas), and its knowledge of health is highly incomplete when compared with dedicated biomedical ontologies. In addition, it is hard to avoid errors such as incorrect facts and categorisations in knowledge graphs; e.g., FoodOn categorises soy milk as a kind of milk, but not as a kind of soy product. Such errors may be inherited from the information source or be caused by the construction procedure. These issues significantly impact the usefulness of knowledge graphs and the reliability of the systems that use them; e.g., the categorisation of soy milk could be dangerous if the knowledge graph were used in a food allergen alert system.Therefore, effective knowledge graph construction and curation is urgently required and will play a critical role in exploiting the full value of knowledge graphs. As there are now many available knowledge resources, one possible approach is to use multiple sources to address both coverage and quality issues, e.g., via integration and cross-checking. For example, integrating HeLis with FoodOn would combine fine-grained categorization of food products (including bananas) with lifestyle knowledge. Moreover, cross-checking FoodOn with HeLis will reveal the problem with soy milk, which is correctly categorized as a soy product in HeLis. Automating the integration of knowledge resources is challenging, but combining semantic and learning-based techniques seems to be a very promising approach, and we have already obtained some encouraging preliminary results in this direction.The proposed research will therefore study a range of semantic and machine learning techniques, and how to combine them to support knowledge graph construction and curation. As well as its application to knowledge graph construction and curation, this research will also contribute to the development of new neural-symbolic theories, paradigms and methods, such as deep semantic embedding for learning representations for expressive knowledge, and knowledge-guided learning for addressing sample shortage problems. These techniques promise to revolutionize many AI and big data technologies.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2207.01328
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
作者:
[Zhuo Chen;Yufen Huang;Jiaoyan Chen;Yuxia Geng;Wen Zhang;Yin Fang;Jeff Z. Pan;Wenting Song;Huajun Chen]
通讯作者:
Zhuo Chen;Yufen Huang;Jiaoyan Chen;Yuxia Geng;Wen Zhang;Yin Fang;Jeff Z. Pan;Wenting Song;Huajun Chen
DOI:
10.1109/jproc.2023.3279374
发表时间:
2021-12
期刊:
Proceedings of the IEEE
影响因子:
20.6
作者:
[Jiaoyan Chen;Yuxia Geng;Zhuo Chen;Jeff Z. Pan;Yuan He;Wen Zhang;Ian Horrocks;Hua-zeng Chen]
通讯作者:
Jiaoyan Chen;Yuxia Geng;Zhuo Chen;Jeff Z. Pan;Yuan He;Wen Zhang;Ian Horrocks;Hua-zeng Chen
DOI:
10.1007/s11280-023-01169-9
发表时间:
2022-02
期刊:
World Wide Web
影响因子:
--
作者:
[Jiaoyan Chen;Yuan He;E. Jiménez-Ruiz;Hang Dong;Ian Horrocks]
通讯作者:
Jiaoyan Chen;Yuan He;E. Jiménez-Ruiz;Hang Dong;Ian Horrocks
DOI:
10.24963/ijcai.2021/597
发表时间:
2021-02
期刊:
Journal of Optics
影响因子:
2.1
作者:
[Jiaoyan Chen;Yuxia Geng;Zhuo Chen;Ian Horrocks;Jeff Z. Pan;Huajun Chen]
通讯作者:
Jiaoyan Chen;Yuxia Geng;Zhuo Chen;Ian Horrocks;Jeff Z. Pan;Huajun Chen
Rewriting the infinite chase
重写无限追逐
DOI:
10.14778/3551793.3551851
发表时间:
2022
期刊:
Proceedings of the VLDB Endowment
影响因子:
2.5
作者:
[Benedikt M]
通讯作者:
Benedikt M
共 6 条
ED3: Enabling analytics over Diverse Distributed Datasources
-
批准号:EP/N014359/1
-
项目类别:Research Grant
-
资助金额:$110.41万
-
财政年份:2016
-
负责人:Ian Horrocks
-
依托单位:
DBOnto: Bridging Databases and Ontologies
-
批准号:EP/L012138/1
-
项目类别:Research Grant
-
资助金额:$161.03万
-
财政年份:2014
-
负责人:Ian Horrocks
-
依托单位:
ExODA: Integrating Description Logics and Database Technologies for Expressive Ontology-Based Data Access
-
批准号:EP/H051511/1
-
项目类别:Research Grant
-
资助金额:$89.77万
-
财政年份:2011
-
负责人:Ian Horrocks
-
依托单位:
ConDOR: Consequence-Driven Ontology Reasoning
-
批准号:EP/G02085X/1
-
项目类别:Research Grant
-
资助金额:$45.83万
-
财政年份:2009
-
负责人:Ian Horrocks
-
依托单位:
HermiT: Reasoning with Large Ontologies
-
批准号:EP/F065841/1
-
项目类别:Research Grant
-
资助金额:$61.21万
-
财政年份:2008
-
负责人:Ian Horrocks
-
依托单位:
LOGO: Logics for Ontologies
-
批准号:EP/C543319/2
-
项目类别:Fellowship
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Ian Horrocks
-
依托单位:
Reasoning Infrastructure for Ontologies and Instances
-
批准号:EP/E03781X/1
-
项目类别:Research Grant
-
资助金额:$73.0万
-
财政年份:2007
-
负责人:Ian Horrocks
-
依托单位:
REOL: Reasoning for Expressive Ontology Languages
-
批准号:EP/C537211/2
-
项目类别:Research Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Ian Horrocks
-
依托单位:
LOGO: Logics for Ontologies
-
批准号:EP/C543319/1
-
项目类别:Fellowship
-
资助金额:$47.04万
-
财政年份:2006
-
负责人:Ian Horrocks
-
依托单位:
海外基金