Taxamatch, an Algorithm for Near ('Fuzzy') Matching of Scientific Names in Taxonomic Databases

Taxamatch, an Algorithm for Near ('Fuzzy') Matching of Scientific Names in Taxonomic Databases
复制标题

DOI:
10.1371/journal.pone.0107510
复制
发表时间:
2014-09-23
期刊:
影响因子:
3.7
通讯作者:
Rees, Tony
Rees, Tony
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Rees, Tony

文献摘要

被引文献

相似文献

生物体科学名称的拼写错误会妨碍生物数据的最佳存储和组织、同一名称不同拼写变体下存储的数据的协调以及用户对分类数据系统的查询的适当回应。本研究提出了一个分析的问题的性质,从第一原则,回顾了一些可用的算法方法,并介绍了Taxamatch,这个信息域的改进的名称匹配解决方案。Taxamatch采用自定义的ModifiedDamerau-Levenshtein Distance算法与语音算法相结合,以及包含一套启发式过滤器的基于规则的方法,与现有的动态编程算法相比,提高了召回率、精确度和执行时间水平n-grams(作为二元组和三元组)和标准编辑距离。虽然完全语音的方法比Taxamatch更快,但它们在召回方面较差,因为许多现实世界的错误本质上是非语音的。Taxamatch的出色性能(如召回率,精度和执行时间)被证明对超过465,000个属名和160万个物种名称的参考数据库,以及对一系列错误类型,目前在属和物种水平的三组样本数据的物种和四个属单独。包括辅助权威匹配组件,其可用于拼写错误的名称和用于在相关联的引用权威不相同的情况下匹配名称。
Misspellings of organism scientific names create barriers to optimal storage and organization of biological data, reconciliation of data stored under different spelling variants of the same name, and appropriate responses from user queries to taxonomic data systems. This study presents an analysis of the nature of the problem from first principles, reviews some available algorithmic approaches, and describes Taxamatch, an improved name matching solution for this information domain. Taxamatch employs a custom Modified Damerau-Levenshtein Distance algorithm in tandem with a phonetic algorithm, together with a rule-based approach incorporating a suite of heuristic filters, to produce improved levels of recall, precision and execution time over the existing dynamic programming algorithms n-grams (as bigrams and trigrams) and standard edit distance. Although entirely phonetic methods are faster than Taxamatch, they are inferior in the area of recall since many real-world errors are non-phonetic in nature. Excellent performance of Taxamatch (as recall, precision and execution time) is demonstrated against a reference database of over 465,000 genus names and 1.6 million species names, as well as against a range of error types as present at both genus and species levels in three sets of sample data for species and four for genera alone. An ancillary authority matching component is included which can be used both for misspelled names and for otherwise matching names where the associated cited authorities are not identical.