Results of AML participation in OAEI 2018

Results of AML participation in OAEI 2018
复制标题

DOI:
--
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
Daniel Faria;Catia Pesquita;B. Balasubramani;Teemu Tervo;David Carriço;Rodrigo Garrilha;Francisco M. Couto;I. Cruz
Daniel Faria;Catia Pesquita;B. Balasubramani;Teemu Tervo;David Carriço;Rodrigo Garrilha;Francisco M. Couto;I. Cruz
中科院分区:
其他
文献类型:
--
作者:
Daniel Faria;Catia Pesquita;B. Balasubramani;Teemu Tervo;David Carriço;Rodrigo Garrilha;Francisco M. Couto;I. Cruz

文献摘要

被引文献

相似文献

AML是一个自动化本体匹配系统,其特点是效率高,可扩展性强,能够整合外部知识。在OAEI 2018中,AML利用这些功能扩展了其处理新赛道的能力。特别的工作是扩展AML以产生复杂的映射,并改进实例匹配方法。AML是今年唯一一个参与所有OAEI轨道的系统,并且在大多数轨道中是表现最好的系统,或者是表现最好的系统之一。1系统介绍1.1状态、目的、总体说明EmphasementMakerLight(AML)是一个基于EmphasementMaker [1,2]设计原则的本体匹配系统,其重点是效率,能够解决大规模的本体匹配问题[7]。它最初的重点是生物医学领域,但它已经不断扩展到解决广泛的本体和实例匹配问题,现在它是一个通用的本体匹配系统。AML主要依赖于词法匹配算法[8],但也包括用于匹配和过滤的结构化算法,以及它自己的逻辑修复算法[10]。它利用外部生物医学本体和WordNet作为背景知识的来源[6]。今年,我们的AML开发主要集中在新的复杂匹配轨道中解决复杂匹配问题。唉,仅仅扩展AML来处理复杂的EDOAL对齐格式就占用了我们大部分的开发时间。当我们最终能够开始开发匹配算法时,很明显,众多类型的EDOAL映射中的每一种都需要自己的专用算法,并且只能为会议数据集中的一些最简单的情况开发算法。我们也无法在OAEI截止日期之前将复杂匹配的代码与主AML代码库完全集成,因此使用不同版本的AML,AMLC参与了复杂匹配跟踪。除了这个版本和主要的AML SEALS版本,我们还通过HOBBIT平台参与了SPIMBENCH和Link Discovery轨道。在SPIMBENCH的情况下,我们参与了主要AML代码库的HOBBIT适应。在链接发现的情况下,我们与OAEI 2017中的情况一样,使用了两个专门版本的AML(分别用于空间和链接任务的AML-空间和AML-链接),这是由于这些匹配任务的独特特性以及HOBBIT数据集中TBox断言的不可用。1.2本节仅描述OAEI 2018新的AML功能。关于AML的匹配策略的更多信息,我们引导读者阅读AML的原始论文[7]以及最近三个版本的OAEI结果出版物[4,5,3]。1.2.1复杂的AML对于复杂的匹配跟踪,我们专注于基于会议本体的挑战。我们开发了基于类似于[9]的模式来识别属性出现限制和属性域限制的策略。通过(1)计算源类和域/范围之间的词汇相似度来检测属性出现限制(2)选择具有高于给定阈值的域/范围相似性的目标属性;(3)对具有相似域的性质建立一个具有比较器和一个非负整数的复映射,增加了对具有类似范围的那些的逆属性限制。通过(1)测量源类和目标类之间的词汇相似度并选择阈值以上的目标类;(2)从源标签中删除匹配的词;(3)将剩余的源串与目标属性匹配并选择阈值以上的目标属性;(4)组成一个复杂映射,该复杂映射被赋予由两个部分相似性(类和属性)加权的分数;(5)选择分数高于阈值的复杂映射。1.2.2主要反洗钱我们仅对该OAEI版本的主要反洗钱代码库进行了一些微小更改。实例匹配在以前的OAEI版本中,AML的实例匹配策略仅依赖于个体的数据属性值和个体之间的关系。今年,由于新的知识图谱轨道,其中个人匹配预计将主要基于他们的注释,AML在其实例匹配库中添加了与它已经用于类和属性匹配的相同的基于词汇的策略。然而,由于在OAEI截止日期之前使用OWL API解析数据集时出现的问题,我们无法正确配置此匹配策略并确保其效率。交互式匹配我们修复了AML交互管理器中的一个错误,该错误导致它在选择和修复步骤之间忘记用户反馈,从而重复一些问题。1.3与去年的情况一样,AML的链接发现提交文件适用于这些特定的任务和数据集,因为它们的特殊性(即没有Tbox)需要专门的提交。在某种程度上,AML的复杂匹配提交也是如此。像往常一样,我们的提交包括翻译的预先计算的字典,以规避微软翻译的查询限制。1.4链接到系统和参数文件AML是一个开源的本体匹配系统,可通过GitHub:https://github.com/AgreementMakerLight获得。
AgreementMakerLight (AML) is a system for automated ontology matching that is characterized by its efficiency, extensibility, and ability to incorporate external knowledge. In OAEI 2018, AML leveraged these features to expand its capabilities to tackle the new tracks. Particular effort was put into extending AML to produce complex mappings, and into improving instance matching approaches. AML was the only system to participate in all OAEI tracks this year, and was the top performing system, or among the top performing systems, in most tracks. 1 Presentation of the System 1.1 State, Purpose, General Statement AgreementMakerLight (AML) is an ontology matching system based on the design principles of AgreementMaker [1, 2] with an added focus on efficiency, to be able to tackle large-scale ontology matching problems [7]. Its initial focus was the biomedical domain, but it has been continually expanded to address a broad range of ontology and instance matching problems, and it is now a general purpose ontology matching system. AML relies primarily on lexical matching algorithms [8], but also includes structural algorithms for both matching and filtering, as well as its own logical repair algorithm [10]. It makes use of external biomedical ontologies and the WordNet as sources of background knowledge [6]. This year, our development of AML was mainly focused on tackling complex matching problems from the new Complex Matching track. Alas, just extending AML to handle the complex EDOAL alignment format took up most of our development time. When we were finally able to start developing matching algorithms, it became clear that each of the numerous types of EDOAL mappings would require its own specialized algorithm, and were only able to develop algorithms for some of the simplest cases, found in the Conference dataset. We were also unable to fully integrate the code for complex matching with the main AML code-base before the OAEI deadline, and thus participated in the Complex Matching track using a different version of AML, AMLC. In addition to this version and the main AML SEALS version, we participated in the SPIMBENCH and Link Discovery tracks via the HOBBIT platform. In the case of SPIMBENCH, we participated with the HOBBIT adaptation of the main AML code-base. In the case of Link Discovery, we participated with two specialized versions of AML (AML-Spatial and AML-Linking for the Spatial and Linking tasks respectively) as had been the case in OAEI 2017, due to the unique characteristics of these matching tasks and to the unavailability of the TBox assertions in the HOBBIT datasets. 1.2 Specific Techniques Used This section describes only the features of AML that are new for the OAEI 2018. For further information on AML’s matching strategy, we direct the reader to AML’s original paper [7] as well as to the OAEI results publications of the last three editions [4, 5, 3]. 1.2.1 Complex AML For the complex matching track, we focused on the challenge based on the conference ontologies. We developed strategies to identify Attribute Occurrence Restrictions and Attribute Domain Restrictions based on patterns similar to [9]. Attribute Occurrence Restrictions were detected by (1) computing the lexical similarities between the source class and the domains/ranges (or superclasses of domains/ranges) of target properties; (2) selecting target properties with domain/range similarity above a given threshold; (3) building a complex mapping with a comparator and a non-negative integer for the properties with similar domain, adding an inverse property restriction for those with similar range. Attribute Domain Restrictions were discovered by (1) measuring the lexical similarity between the source class and target classes and selecting target classes above a threshold; (2) removing the matched words from source labels; (3) matching the remaining source strings to target properties and selecting target properties above a threshold; (4) composing a complex mapping which is given a score weighted by the two partial similarities (class and property); (5) selecting complex mappings with scores above a threshold. 1.2.2 Main AML We made only a few minor changes to the main AML code-base for this OAEI edition. Instance Matching In previous OAEI editions, AML’s matching strategy for instance matching relied only on Data Property values of individuals and on the relations between individuals. This year, due to the new Knowledge Graph track in which individual matching is expected to be mainly based on their annotations, AML added to its instance matching arsenal the same lexical-based strategy it was already using for class and property matching. However, due to problems in parsing the datasets with the OWL API before the OAEI deadline, we were unable to properly configure this matching strategy and ensure its efficiency. Interactive Matching We fixed a bug in AML’s interaction manager that was causing it to forget user feedback between the selection and repair steps and thus repeat some questions. 1.3 Adaptations made for the evaluation As was the case last year, the Link Discovery submissions of AML are adapted to these particular tasks and datasets, as their specificities (namely the absence of a Tbox) demand a dedicated submission. The same is also true to some extent of AML’s Complex Matching submission. As usual, our submission included precomputed dictionaries with translations, to circumvent Microsoftr Translator’s query limit. 1.4 Link to the system and parameters file AML is an open source ontology matching system and is available through GitHub: https://github.com/AgreementMakerLight.