Input-output HMM's for sequence processing

Input-output HMM's for sequence processing
复制标题

DOI:
10.1109/72.536317
复制
发表时间:
1996-09-01
影响因子:
--
通讯作者:
Frasconi, P
Frasconi, P
中科院分区:
其他
文献类型:
--
作者:
Bengio, Y;Frasconi, P

文献摘要

被引文献

相似文献

我们考虑了序列处理的问题,并提出了一种基于离散状态模型的解决方案,以表示过去的上下文。我们引入了一个循环连接主义架构,该架构具有模块化结构,将子网络与每个状态关联起来,该模型具有我们称为输入输出隐马尔可夫模型(IOHMM)的统计解释。它可以通过估计最大化(EM)或广义EM (GEM)算法进行训练,将状态轨迹视为缺失数据,从而解耦了时间信用分配和实际参数估计。该模型与隐马尔可夫模型(HMM)相似,但允许我们将输入序列映射到输出序列,使用与递归神经网络相同的处理风格。IOHMM使用比HMM更具判别性的学习范式进行训练,同时潜在地利用了EM算法的优势。我们在一个基准问题上证明了IOHMM非常适合于解决语法推理问题,并给出了7种Tomita语法的实验结果,表明这些自适应模型可以获得很好的泛化。
We consider problems of sequence processing and propose a solution based on a discrete-state model in order to represent past context, We introduce a recurrent connectionist architecture having a modular structure that associates a subnetwork to each state, The model has a statistical interpretation we call input-output hidden Markov model (IOHMM). It can be trained by the estimation-maximization (EM) or generalized EM (GEM) algorithms, considering state trajectories as missing data, which decouples temporal credit assignment and actual parameter estimation, The model presents similarities to hidden Markov models (HMM's), but allows us to map input sequences to output sequences, using the same processing style as recurrent neural networks. IOHMM's are trained using a more discriminant learning paradigm than HMM's, while potentially taking advantage of the EM algorithm. We demonstrate that IOHMM's are well suited for solving grammatical inference problems on a benchmark problem, Experimental results are presented for the seven Tomita grammars, showing that these adaptive models can attain excellent generalization.