PhyloParser: A Hybrid Algorithm for Extracting Phylogenies from Dendrograms

PhyloParser: A Hybrid Algorithm for Extracting Phylogenies from Dendrograms
复制标题

PhyloParser:一种从树状图中提取系统发育的混合算法

DOI:
10.1109/icdar.2017.180
复制
发表时间:
2017
期刊:
2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)
影响因子:
--
通讯作者:
Bill Howe
Bill Howe
中科院分区:
--
文献类型:
--
作者:
Po;Sean T. Yang;Jevin D. West;Bill Howe

文献摘要

参考文献

被引文献

相似文献

我们认为一种新的方法来提取信息的树状图在生物文献中代表系统发育树。现有的算法方法来提取这些关系依赖于跟踪树的轮廓,是非常敏感的图像质量问题,但手动方法需要大量的人力,不能在规模上使用。我们介绍PhyloParser,一个完全自动化的,端到端的系统,用于自动提取物种的关系,从系统发育树图使用多模式的方法来消化不同的树风格。我们的方法自动识别系统发育树的数字在科学文献中,提取树结构的关键组成部分,重建树,恢复物种关系。我们使用多种方法来提取具有高召回率的树组件,然后通过对这些组件如何组合在一起应用拓扑分析来过滤误报。我们提出了一个真实世界的数据集上的评估,定量和定性地证明我们的方法的有效性。我们的分类器实现了89%的召回率和99%的准确率,相对于以前的方法具有较低的平均错误率。我们的目标是使用PhyloParser建立一个链接的,开放的,全面的系统发育信息数据库,涵盖历史文献以及当前的数据,然后使用此资源来确定生物学文献中的分歧和覆盖率低的领域。
We consider a new approach to extracting information from dendrograms in the biological literature representing phylogenetic trees. Existing algorithmic approaches to extract these relationships rely on tracing tree contours and are very sensitive to image quality issues, but manual approaches require significant human effort and cannot be used at scale. We introduce PhyloParser, a fully automated, end-to-end system for automatically extracting species relationships from phylogenetic tree diagrams using a multi-modal approach to digest diverse tree styles. Our approach automatically identifies phylogenetic tree figures in the scientific literature, extracts the key components of tree structure, reconstructs the tree, and recovers the species relationships. We use multiple methods to extract tree components with high recall, then filter false positives by applying topological heuristics about how these components fit together. We present an evaluation on a real-world dataset to quantitatively and qualitatively demonstrate the efficacy of our approach. Our classifier achieves 89% recall and 99% precision, with a low average error rate relative to previous approaches. We aim to use PhyloParser to build a linked, open, comprehensive database of phylogenetic information that covers the historical literature as well as current data, and then use this resource to identify areas of disagreement and poor coverage in the biological literature.
从数千篇论文的数据中挖掘出机器编译的微生物超级树
DOI: 10.3897/rio.3.e13589
发表时间: 2017
影响因子: --
作者:
Mounce R
通讯作者: Mounce R
DOI: 10.1093/molbev/msp296
发表时间: 2010-04-01
影响因子: 10.7
作者:
Hey, Jody
通讯作者: Hey, Jody