Using lexicography to characterise relations between species mentions in the biodiversity literature

Using lexicography to characterise relations between species mentions in the biodiversity literature
复制标题

DOI:
10.1145/3322905.3322918
复制
发表时间:
2019-05
期刊:
Proceedings of the 3rd International Conference on Digital Access to Textual Cultural Heritage
影响因子:
--
通讯作者:
Sandra Young
Sandra Young
中科院分区:
其他
文献类型:
--
作者:
Sandra Young

文献摘要

相似文献

生物多样性文献是世界上记录遗产的历史最悠久的例子之一。今天,为了遗产和研究的目的,有许多努力来收集和整合文献,以确保获得信息。在这些努力中,本体越来越多地被用作知识表示工具。然而,使用本体框架来表示生物分类的有效性受到了质疑。生物分类学使用科学命名法为所描述的物种命名。虽然命名法是一个有用的分类工具,但由于其同义、同源和流动的性质,它也可能成为混淆的来源。尽管如此,在文献中使用的科学术语没有经验评价已经进行过。基于语料库的分析已经被用于自动本体提取,本研究探讨了应用最近开发的词典编纂技术的问题,以提供评估的文献中的经验数据的可能性,并作为与现有的本体比较。本文重点介绍了如何从文献中提取结构进行这些比较的研究的工作流程,参数和初步结果。它使用的语料库分析技术,可视化和过滤方法的操作,这样做,并评估潜在的分类和消歧质量的结果图为未来的工作。初步的结果看频率和显着性的影响时,过滤的图形,这表明,这些过滤器参数可以用于不同的目的,在揭示有机体之间的关系提到。
The biodiversity literature is one of the longest-standing examples of recording heritage in the world. Today there are many efforts to standardise and integrate the literature to ensure access to the information, both for heritage and research purposes. Ontologies are increasingly being turned to as knowledge representation tools in these efforts. However, the validity of using ontological frameworks to represent biological taxonomies has been questioned. Biological taxonomies use the scientific nomenclature to assign names to described species. While the nomenclature is a useful classification tool, it can also be a source of confusion because of its synonymous, homonymous and fluid nature. Despite this, no empirical evaluation of scientific nomenclature use in the literature has ever been performed. Corpus-based analysis is already used in automatic ontology extraction, and this study explores the possibility of applying recently developed lexicography techniques to the problem to provide an evaluation of the empirical data in the literature, and serve as a comparison with existing ontologies. This paper focuses on the work flow, parameters and preliminary findings of the research investigating how to extract structures from the literature to perform these comparisons. It uses the manipulation of corpus analysis techniques, visualisation and filtering methods to do so and evaluates potential classification and disambiguation qualities of the resulting graphs for future work. Preliminary results look at the effects of frequency and salience when filtering the graphs, which indicate that these filter parameters could be used for different purposes in revealing relationships between organism mentions.