Analysis of Book Documents’ Table of Content Based on Clustering

Analysis of Book Documents’ Table of Content Based on Clustering
复制标题

基于聚类的图书文献目录分析

DOI:
--
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
Yimin Chu
Yimin Chu
中科院分区:
--
文献类型:
--
作者:
Liangcai Gao;Zhi Tang;Xiaofan Lin;Xin Tao;Yimin Chu

文献摘要

被引文献

相似文献

目录(TOC)识别近年来引起了人们的极大关注。在回顾了现有目录识别方法的优缺点后,我们观察到图书文档是具有固有的本地格式一致性的多页文档。基于这一发现,我们提出了一种基于聚类的TOC自动分析方法。此方法首先检测TOC页面中的装饰元素。然后通过聚类来学习TOC页面中使用的布局模型。最后,在该模型的指导下,生成TOC条目并提取其层次结构。更具体地说,在该方法中考虑了虚线。实验结果表明,该方法具有较高的准确率和效率。此外,该方法已成功应用于一个商品化的电子书制作软件包中。
Table of contents (TOC) recognition has attracted a great deal of attention in recent years. After reviewing the merits and drawbacks of the existing TOC recognition methods, we have observed that book documents are multi-page documents with intrinsic local format consistency. Based on this finding we introduce an automatic TOC analysis method through clustering. This method first detects the decorative elements in TOC pages. Then it learns a layout model used in the TOC pages through clustering. Finally, it generates TOC entries and extracts their hierarchical structure under the guidance of the model. More specifically, broken lines are taken into account in the method. Experimental results show that this method achieves high accuracy and efficiency. In addition, this method has been successfully applied in a commercial E-book production software package.