Distilling an Ensemble of Greedy Dependency Parsers into One MST Parser

Distilling an Ensemble of Greedy Dependency Parsers into One MST Parser
复制标题

DOI:
10.18653/v1/d16-1180
复制
发表时间:
2016-09
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
A. Kuncoro;Miguel Ballesteros;Lingpeng Kong;Chris Dyer;Noah A. Smith
A. Kuncoro;Miguel Ballesteros;Lingpeng Kong;Chris Dyer;Noah A. Smith
中科院分区:
其他
文献类型:
--
作者:
A. Kuncoro;Miguel Ballesteros;Lingpeng Kong;Chris Dyer;Noah A. Smith

文献摘要

被引文献

相似文献

我们引入了两个一阶图依赖解析器,达到了一个新的技术水平。第一种是共识解析器,由具有不同随机初始化的独立训练的贪婪LSTM转换解析器集成而成。我们将这种方法称为最小贝叶斯风险译码(在汉明成本下),并认为在集合内较弱的共识是困难或歧义的有用信号。第二个解析器是将集合“升华”成单一模型。我们使用具有新成本的结构化铰链损失目标来训练蒸馏解析器,该成本结合了对每个可能的附件的集成不确定性估计,从而避免了将标准蒸馏目标应用于具有结构化输出的问题所需的棘手的交叉熵计算。一阶蒸馏解析器在英语、汉语和德语方面达到或超过了最先进的水平。
We introduce two first-order graph-based dependency parsers achieving a new state of the art. The first is a consensus parser built from an ensemble of independently trained greedy LSTM transition-based parsers with different random initializations. We cast this approach as minimum Bayes risk decoding (under the Hamming cost) and argue that weaker consensus within the ensemble is a useful signal of difficulty or ambiguity. The second parser is a "distillation" of the ensemble into a single model. We train the distillation parser using a structured hinge loss objective with a novel cost that incorporates ensemble uncertainty estimates for each possible attachment, thereby avoiding the intractable cross-entropy computations required by applying standard distillation objectives to problems with structured outputs. The first-order distillation parser matches or surpasses the state of the art on English, Chinese, and German.