A Numerical Study of the Bottom-Up and Top-Down Inference Processes in And-Or Graphs

A Numerical Study of the Bottom-Up and Top-Down Inference Processes in And-Or Graphs
复制标题

与或图自下而上和自上而下推理过程的数值研究

DOI:
10.1007/s11263-010-0346-6
复制
发表时间:
2011-06-01
影响因子:
19.5
通讯作者:
Zhu, Song-Chun
Zhu, Song-Chun
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wu, Tianfu;Zhu, Song-Chun

文献摘要

被引文献

相似文献

本文以“与或”图为例,对分层模型中自下而上和自上而下的推理过程进行了数值研究。为递归定义的与或图中的每个节点 A 确定三个推理过程,其中嵌入了随机上下文敏感图像语法:alpha(A) 过程直接基于图像特征检测节点 A,beta(A) 过程通过自下而上绑定其子节点来计算节点 A,gamma(A) 过程从其父节点自上而下预测节点 A。所有三个进程都以互补的方式从图像中计算节点 A。我们数值研究的目的是探索每个流程贡献多少信息以及如何整合这些流程以提高绩效。我们使用贝叶斯框架下制定的“与或”图在对象解析任务中研究它们。首先,我们通过阻塞其他两个进程来分别隔离和训练 alpha(A)、beta(A) 和 gamma(A) 进程。然后,根据每个过程的辨别力并与各自的人类表现进行比较,单独评估每个过程的信息贡献。其次,我们显式地集成三个过程以进行鲁棒推理以提高性能,并提出一种用于对象解析的贪婪追踪算法。在实验中,我们选择了两个层次案例研究:一个是中低级视觉中的路口和矩形,另一个是高级视觉中的人脸。我们观察到:(i) alpha(A)、beta(A) 和 gamma(A) 过程的有效性取决于尺度和遮挡条件,(ii) alpha(face) 过程比面部组件的 alpha 过程更强,而 beta(junctions) 和 beta(矩形) 比它们的 alpha 过程更好,(iii) 三个过程的集成提高了 ROC 比较中的性能。
This paper presents a numerical study of the bottom-up and top-down inference processes in hierarchical models using the And-Or graph as an example. Three inference processes are identified for each node A in a recursively defined And-Or graph in which stochastic context sensitive image grammar is embedded: the alpha(A) process detects node A directly based on image features, the beta(A) process computes node A by binding its child node(s) bottom-up and the gamma(A) process predicts node A top-down from its parent node(s). All the three processes contribute to computing node A from images in complementary ways. The objective of our numerical study is to explore how much information each process contributes and how these processes should be integrated to improve performance. We study them in the task of object parsing using And-Or graph formulated under the Bayesian framework. Firstly, we isolate and train the alpha(A), beta(A) and gamma(A) processes separately by blocking the other two processes. Then, information contributions of each process are evaluated individually based on their discriminative power, compared with their respective human performance. Secondly, we integrate the three processes explicitly for robust inference to improve performance and propose a greedy pursuit algorithm for object parsing. In experiments, we choose two hierarchical case studies: one is junctions and rectangles in low-to-middle-level vision and the other is human faces in high-level vision. We observe that (i) the effectiveness of the alpha(A), beta(A) and gamma(A) processes depends on the scale and occlusion conditions, (ii) the alpha(face) process is stronger than the alpha processes of facial components, while beta(junctions) and beta(rectangle) work much better than their alpha processes, and (iii) the integration of the three processes improves performance in ROC comparisons.