DyGen: Learning from Noisy Labels via Dynamics-Enhanced Generative Modeling

DyGen: Learning from Noisy Labels via Dynamics-Enhanced Generative Modeling
复制标题

DOI:
10.1145/3580305.3599318
复制
发表时间:
2023-05
期刊:
Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Yuchen Zhuang;Yue Yu;Lingkai Kong;Xiang Chen;Chao Zhang
Yuchen Zhuang;Yue Yu;Lingkai Kong;Xiang Chen;Chao Zhang
中科院分区:
其他
文献类型:
--
作者:
Yuchen Zhuang;Yue Yu;Lingkai Kong;Xiang Chen;Chao Zhang

文献摘要

被引文献

相似文献

从嘈杂的标签中学习是一个挑战,它在许多现实世界应用程序中都会出现,在这些应用程序中,培训数据可能包含不正确或损坏的标签。当带有嘈杂标签的微调语言模型时,模型可以轻松地过度贴合标签噪声,从而导致性能下降。从嘈杂标签中学习的大多数现有方法都使用静态输入特征进行降级,但是这些方法受到可以在真实标签分布上提供的信息的限制,并且可能导致偏见或不正确的预测。在这项工作中,我们提出了动态增强生成模型(DYGEN),该模型在语言模型的微调过程中使用嵌入式空间中的动态模式来改善嘈杂的标签预测。 Dygen使用各种自动编码框架来推断嘈杂标签和训练动力学的真实标签的后验分布。此外,使用共同指导机制来最大程度地减少潜在的嘈杂标签和先验的影响。与先前的最新技术相比,DYGEN在两个合成噪声数据集的平均准确性提高了3.10%,三个现实世界噪声数据集的平均准确性提高了1.48%。广泛的实验和分析显示了每个成分在dygen中的有效性。我们的代码可在GitHub上可重复可重复。
Learning from noisy labels is a challenge that arises in many real-world applications where training data can contain incorrect or corrupted labels. When fine-tuning language models with noisy labels, models can easily overfit the label noise, leading to decreased performance. Most existing methods for learning from noisy labels use static input features for denoising, but these methods are limited by the information they can provide on true label distributions and can result in biased or incorrect predictions. In this work, we propose the Dynamics-Enhanced Generative Model (DyGen), which uses dynamic patterns in the embedding space during the fine-tuning process of language models to improve noisy label predictions. DyGen uses the variational auto-encoding framework to infer the posterior distributions of true labels from noisy labels and training dynamics. Additionally, a co-regularization mechanism is used to minimize the impact of potentially noisy labels and priors. DyGen demonstrates an average accuracy improvement of 3.10% on two synthetic noise datasets and 1.48% on three real-world noise datasets compared to the previous state-of-the-art. Extensive experiments and analyses show the effectiveness of each component in DyGen. Our code is available for reproducibility on GitHub.