AI and Cognitive Computing for Reasoning about Big Data and Knowledge Graphs with Application to the Oil and Gas Industry
AI and Cognitive Computing for Reasoning about Big Data and Knowledge Graphs with Application to the Oil and Gas Industry
批准号:
2370505
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
简要描述:该博士项目的主要目标是与BP合作,深入了解如何将最新的人工智能和认知计算技术用于大数据推理。该项目的一个应用是在石油和天然气行业,支持和改进该行业的核心业务流程和决策。为了实现这一目标,该学生计划的重要贡献是开发大数据规则学习器和推理器,为存在规则设计学习器,并系统评估其对BP重要应用的影响。为了学生、英国石油公司和牛津大学的共同利益,并最大限度地发挥协同作用,该学生将加入VADA“增值数据系统”项目,该项目的目标之一是基于Datalog语言系列的经验,开发一个通用的推理系统。与EPSRC的战略和研究领域保持一致:该项目属于EPSRC信息和通信技术(ICT)主题,以及以下研究领域:人工智能技术,数据库和信息系统。新颖性和研究方法论:声明性规则,如Prolog和Datalog规则是表达专业知识的常见形式,在许多系统中使用。由于开发这样的规则是耗时的,并且需要很少的专家知识,因此开发用于学习这些规则的算法是必不可少的。该项目解决了存在规则的学习问题,存在规则在许多用例中都有应用,如知识图、语义网和Web数据提取。特别是,我们专注于开发存在规则的进化学习算法。我们定义了规则学习设置,并回顾了学习规则的主要方法,如自顶向下、自底向上和神经方法。我们回顾了现有的规则学习的进化方法,讨论了不同的遗传编码模式、初始群体创建方法、进化算子和评估适应度函数。此外,从更广泛的角度来看,我们探索了逻辑推理引擎和机器学习方法之间的四种交互模型。最后,我们概述了存在规则学习提出的研究问题的答案,并展示了存在规则学习在原油腐蚀性、知识图规则挖掘和问答数据集方面的应用。本项目重点研究归纳逻辑规划(ILP)问题的进化算法(EA,也称为遗传算法,GA)。ea是一个受生物学启发的搜索算法家族,它优化最有前途的初步解决方案,同时探索广泛的搜索空间。特别是,在我们的设置中,原子、部分规则或参数化规则可以被视为染色体,通过突变、交叉和选择的操作,可以得到新的染色体群体。在执行这些操作时,计算一个称为适应度函数的质量度量来判断获得的新一代染色体是否适合继续搜索。这种进化算法通常不执行穷举搜索,同时也不太可能陷入局部最优。此外,它们是灵活的,因为它们可能不需要在规则的形状上强加模板,因为这是其他ILP方法的典型情况。参与的公司和合作者:这是EPSRC与英国石油公司合作的工业案例学生项目。
英文摘要
Brief Description: The broad aim of this doctoral project is to gain, in cooperation with BP, a deep understanding of how the latest AI and cognitive computing technologies can be used for reasoning over big data. One of the application of this project is in the oil and gas industry - supporting and improving core business processes and decision making in this sector. Towards this aim, the significant contributions of the student are planned to be the development of a rule learner and reasoner for big data, specific designed of the learner for existential rules, and systematic evaluation its impact on applications important to BP.To the mutual benefit of the student, BP, and Oxford University, and to maximise synergies, the studentship will be attached to the VADA "Value Added Data Systems" project which - as one of its goals - aims to develop a general-purpose reasoning system, building on the experience with the Datalog family of languages.Alignment to EPSRC's Strategies and Research Areas: This project falls within the EPSRC Information and communication technologies (ICT) theme, and the following research areas: Artificial intelligence technologies, Databases, and Information systems.Novelty and Research Methodology: Declarative rules such as Prolog and Datalog rules are common formalisms to express expert knowledge and are used in a number of systems. Since developing such rules is time-consuming and requires scarce expert knowledge, it is essential to develop algorithms for learning such rules. This project addresses the problem of learning existential rules, which found applications in many uses cases such as Knowledge Graphs, the Semantic Web and Web Data Extraction. In particular, we concentrate on developing evolutionary learning algorithms for existential rules. We define the rule learning setting and review the main approaches to learning rules, such as top-down, bottom-up, and neural methods. We review existing evolutionary approaches to rule learning, discuss different genetic encoding schema, initial population creation methods, evolution operators, and evaluation fitness functions. In addition, from a wider view, we explore four interaction models between logical reasoning engines and Machine Learning approaches. Last but not least, we outline the answers to the proposed research questions for existential rule learning with promising experimental results, and exhibit applications in crude corrosivity, knowledge graph rule mining and question answering data sets. This project focuses on studying Evolutionary Algorithms (EA, also known as Genetic Algorithm, GA) for the Inductive Logic Programming (ILP) problem. EAs are a family of biology-inspired search algorithms that optimize for the most promising preliminary solutions, while exploring a wide search space at the same time. In particular, in our setting, atoms, partial or parameterized rules can be treated as chromosomes, from which the new population of chromosomes can be derived via the operations of mutation, crossover and selection. While performing these operations, a quality measure, called fitness function is computed to judge whether an obtained new generation of chromosomes is fit for continuing the search. Such evolutionary algorithms typically do not perform exhaustive search and at the same time are less likely to fall into local optima. In addition, they are flexible in that they might not require imposed template on the shape of rules, as it is typically the case in other approaches to ILP. Companies and Collaborators Involved: This is an EPSRC Industrial CASE studentship project in collaboration with BP.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1609/aaai.v34i05.6216
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
作者:
[Yuhang Song;Jianyi Wang;Thomas Lukasiewicz;Zhenghua Xu;Mai Xu;Zihan Ding;Lianlong Wu]
通讯作者:
Yuhang Song;Jianyi Wang;Thomas Lukasiewicz;Zhenghua Xu;Mai Xu;Zihan Ding;Lianlong Wu
DOI:
10.24963/ijcai.2019/928
发表时间:
2019
期刊:
影响因子:
--
作者:
[Wu L]
通讯作者:
Wu L
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Wu L]
通讯作者:
Wu L
Rule Learning over Knowledge Graphs with Genetic Logic Programming (Extended Abstract)
使用遗传逻辑编程进行知识图的规则学习(扩展摘要)
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Wu L]
通讯作者:
Wu L
Democratise Financial Knowledge Graph Construction By Mining Massive Brokerage Research Report
挖掘海量券商研究报告,民主化金融知识图谱建设
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Cheng Z]
通讯作者:
Cheng Z
共 8 条
海外基金