Sparse Conditional Hidden Markov Model for Weakly Supervised Named Entity Recognition

Sparse Conditional Hidden Markov Model for Weakly Supervised Named Entity Recognition
复制标题

DOI:
10.1145/3534678.3539247
复制
发表时间:
2022-05
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Yinghao Li;Le Song;Chao Zhang
Yinghao Li;Le Song;Chao Zhang
中科院分区:
其他
文献类型:
--
作者:
Yinghao Li;Le Song;Chao Zhang

文献摘要

相似文献

弱监督命名实体识别方法训练标签模型,以聚合多个噪声标签函数(LF)的标记注释,而不会看到任何手动注释的标签。为了更好地工作,标签模型需要根据上下文识别和强调表现良好的LF,同时降低表现不佳的权重。然而,由于缺乏地面实况,评估LF是具有挑战性的。为了解决这个问题,我们提出了稀疏条件隐马尔可夫模型(稀疏CHMM)。与其他基于HMM的方法不同,Sparse-CHMM侧重于估计其对角元素,这些元素被认为是LF的可靠性得分。然后,稀疏分数扩展到具有预定义扩展函数的成熟的发射矩阵。我们还使用加权XOR分数来增强发射,该分数跟踪LF观察到不正确实体的概率。稀疏CHMM通过无监督学习进行优化,采用三阶段训练管道,降低了训练难度,防止模型陷入局部最优。与Wrench基准中的基线相比,Sparse-CHMM在五个综合数据集上实现了平均3.01的F1得分改善。实验表明,稀疏CHMM的每个组成部分是有效的,估计的LF可靠性与真实LF F1分数强相关。
Weakly supervised named entity recognition methods train label models to aggregate the token annotations of multiple noisy labeling functions (LFs) without seeing any manually annotated labels. To work well, the label model needs to contextually identify and emphasize well-performed LFs while down-weighting the under-performers. However, evaluating the LFs is challenging due to the lack of ground truths. To address this issue, we propose the sparse conditional hidden Markov model (Sparse-CHMM). Instead of predicting the entire emission matrix as other HMM-based methods, Sparse-CHMM focuses on estimating its diagonal elements, which are considered as the reliability scores of the LFs. The sparse scores are then expanded to the full-fledged emission matrix with pre-defined expansion functions. We also augment the emission with weighted XOR scores, which track the probabilities of an LF observing incorrect entities. Sparse-CHMM is optimized through unsupervised learning with a three-stage training pipeline that reduces the training difficulty and prevents the model from falling into local optima. Compared with the baselines in the Wrench benchmark, Sparse-CHMM achieves a 3.01 average F1 score improvement on five comprehensive datasets. Experiments show that each component of Sparse-CHMM is effective, and the estimated LF reliabilities strongly correlate with true LF F1 scores.