DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4

DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4
复制标题

DOI:
10.18653/v1/2023.emnlp-main.519
复制
发表时间:
2023-05
期刊:
--
影响因子:
--
通讯作者:
Ye Hu;Kaiqiang Song;Sangwoo Cho;Xiaoyang Wang;H. Foroosh;Fei Liu
Ye Hu;Kaiqiang Song;Sangwoo Cho;Xiaoyang Wang;H. Foroosh;Fei Liu
中科院分区:
其他
文献类型:
--
作者:
Ye Hu;Kaiqiang Song;Sangwoo Cho;Xiaoyang Wang;H. Foroosh;Fei Liu

文献摘要

相似文献

人类的偏好判断在指导大型语言模型(LLM)产生与人类价值观一致的输出方面起着关键作用。人工评估还用于摘要任务,以比较来自不同系统的输出,补充现有的自动指标。然而,尽管它们意义重大,但探索这些成对比较或按$k$比较的研究有限。产出长度、信息性、流畅性和事实一致性等因素的集体影响和相对重要性仍未得到很好的理解。目前也不清楚是否还有其他隐藏因素影响人类的判断。在本文中,我们对OpenAI发布的一组成对的人类判断进行了深入的研究。利用Bradley-Terry-Luce(BTL)模型,我们揭示了这些人类判断中嵌入的内在偏好。我们发现,最受欢迎的因素因任务和类型的不同而不同,而最不受欢迎的因素往往是一致的,例如,输出过于简短,包含过多的离焦内容或幻觉事实。我们的发现对构建人类偏好评估中的平衡数据集具有重要意义,这是塑造未来LLM行为的关键一步。
Human preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values. Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing existing automatic metrics. Despite their significance, however, there has been limited research probing these pairwise or $k$-wise comparisons. The collective impact and relative importance of factors such as output length, informativeness, fluency, and factual consistency are still not well understood. It is also unclear if there are other hidden factors influencing human judgments. In this paper, we conduct an in-depth examination of a collection of pairwise human judgments released by OpenAI. Utilizing the Bradley-Terry-Luce (BTL) model, we reveal the inherent preferences embedded in these human judgments. We find that the most favored factors vary across tasks and genres, whereas the least favored factors tend to be consistent, e.g., outputs are too brief, contain excessive off-focus content or hallucinated facts. Our findings have implications on the construction of balanced datasets in human preference evaluations, which is a crucial step in shaping the behaviors of future LLMs.