Uncertainty-Aware Bootstrap Learning for Joint Extraction on Distantly-Supervised Data

Uncertainty-Aware Bootstrap Learning for Joint Extraction on Distantly-Supervised Data
复制标题

DOI:
10.48550/arxiv.2305.03827
复制
发表时间:
2023-05
期刊:
--
影响因子:
--
通讯作者:
Yufei Li;Xiao Yu;Yanchi Liu;Haifeng Chen;Cong Liu
Yufei Li;Xiao Yu;Yanchi Liu;Haifeng Chen;Cong Liu
中科院分区:
其他
文献类型:
--
作者:
Yufei Li;Xiao Yu;Yanchi Liu;Haifeng Chen;Cong Liu

文献摘要

相似文献

在处理具有模糊或噪声标签的远程监督数据时,联合提取实体对及其关系是一项挑战。为了减轻这种影响,我们提出了不确定性感知引导学习,其动机是直觉,实例的不确定性越高,模型置信度与地面事实不一致的可能性就越大。具体来说,我们首先探索实例级数据的不确定性以创建初始的高置信度示例。这样的子集可以过滤噪声实例并促进模型在早期快速收敛。在引导学习期间,我们提出自集成作为正则化器,以减轻噪声标签产生的模型间不确定性。我们进一步定义联合标记概率的概率方差来估计内部模型参数不确定性,用于为下一次迭代选择和构建新的可靠训练实例。两个大型数据集上的实验结果表明,我们的方法优于现有的强基线和相关方法。
Jointly extracting entity pairs and their relations is challenging when working on distantly-supervised data with ambiguous or noisy labels.To mitigate such impact, we propose uncertainty-aware bootstrap learning, which is motivated by the intuition that the higher uncertainty of an instance, the more likely the model confidence is inconsistent with the ground truths.Specifically, we first explore instance-level data uncertainty to create an initial high-confident examples. Such subset serves as filtering noisy instances and facilitating the model to converge fast at the early stage.During bootstrap learning, we propose self-ensembling as a regularizer to alleviate inter-model uncertainty produced by noisy labels. We further define probability variance of joint tagging probabilities to estimate inner-model parametric uncertainty, which is used to select and build up new reliable training instances for the next iteration.Experimental results on two large datasets reveal that our approach outperforms existing strong baselines and related methods.