Shape quantization and recognition with randomized trees

Shape quantization and recognition with randomized trees
复制标题

DOI:
10.1162/neco.1997.9.7.1545
复制
发表时间:
1997-10-01
期刊:
影响因子:
2.9
通讯作者:
Geman, D
Geman, D
中科院分区:
计算机科学4区
文献类型:
--
作者:
Amit, Y;Geman, D

文献摘要

被引文献

相似文献

我们探索了一种新的方法,形状识别的基础上,一个几乎无限的家庭的二进制功能(查询)的图像数据,旨在适应先验信息的形状不变性和规律性。每个查询对应于几个本地地形代码(或标签)的空间排列,这些代码本身过于原始和常见,无法提供有关形状的信息。所有的鉴别能力都来自标签之间的相对角度和距离。查询的重要属性是对应于增加的结构和复杂性的自然偏序;半不变性,意味着给定类的大多数形状将以相同的方式回答两个在顺序上连续的查询;和稳定性,因为查询不是基于可区分的点和子结构。没有基于完整特征集的分类器可以被评估,并且不可能先验地确定哪些布置是信息性的。我们的方法是选择信息的功能,并建立树分类器在同一时间通过归纳学习。实际上,每棵树都提供了对完整后验的近似,其中所选择的特征取决于所遍历的分支。由于查询的数量和性质,基于固定长度特征向量的标准决策树构造是不可行的。相反,我们在每个节点上只接受一个小的随机查询样本,限制其复杂性随树深度增加,并生长多棵树。终端节点标记的形状类的相应的后验分布的估计。该方法通过将图像发送到每棵树下并聚集得到的分布来进行分类,并应用于手写数字和300个LAT(E)X符号的综合线性和非线性变形的分类。国家的最先进的错误率达到国家标准和技术研究所的数字数据库。LAT(E)X符号实验的主要目标是分析不变性,泛化误差和相关问题,并在此背景下与人工神经网络方法进行比较。
We explore a new approach to shape recognition based on a virtually infinite family of binary features (queries) of the image data, designed to accommodate prior information about shape invariance and regularity. Each query corresponds to a spatial arrangement of several local topographic codes (or tags), which are in themselves too primitive and common to be informative about shape. All the discriminating power derives from relative angles and distances among the tags. The important attributes of the queries are a natural partial ordering corresponding to increasing structure and complexity; semi-invariance, meaning that most shapes of a given class will answer the same way to two queries that are successive in the ordering; and stability, since the queries are not based an distinguished points and substructures.No classifier based on the full feature set can be evaluated, and it is impossible to determine a priori which arrangements are informative. Our approach is to select informative features and build tree classifiers at the same time by inductive learning. In effect, each tree provides an approximation to the full posterior where the features chosen depend on the branch that is traversed. Due to the number and nature of the queries, standard decision tree construction based on a fixed-length feature vector is not feasible. Instead we entertain only a small random sample of queries at each node, constrain their complexity to increase with tree depth, and grow multiple trees. The terminal nodes are labeled by estimates of the corresponding posterior distribution over shape classes. An image is classified by sending it down every tree and aggregating the resulting distributions.The method is applied to classifying handwritten digits and synthetic linear and nonlinear deformations of three hundred LAT(E)X symbols. State-of-the-art error rates are achieved on the National Institute of Standards and Technology database of digits. The principal goal of the experiments on LAT(E)X symbols is to analyze invariance, generalization error and related issues, and a comparison with artificial neural networks methods is presented in this context.