Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the “What” and Focus on the “How”!

Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the “What” and Focus on the “How”!
复制标题

DOI:
10.18653/v1/2023.acl-long.530
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Chen Chen-Chen;Dylan Walker;Venkatesh Saligrama
Chen Chen-Chen;Dylan Walker;Venkatesh Saligrama
中科院分区:
其他
文献类型:
--
作者:
Chen Chen-Chen;Dylan Walker;Venkatesh Saligrama

文献摘要

相似文献

提出了一种新的用于政治意识形态预测(PIP)的监督学习方法,该方法能够预测非分布的输入。这个问题的起因是,手动数据标记昂贵,而自我报告的标记往往很少,并表现出显著的选择偏差。我们提出了一种新的统计模型,将文档嵌入分解为两个向量的线性叠加:与意识形态无关的潜在中性上下文向量和与意识形态对齐的潜在位置向量。我们训练一个端到端模型,该模型将中间的上下文和位置向量作为输出。在部署时,我们的模型通过独占利用预测的位置向量来预测输入文档的标签。在两个基准数据集上,我们的模型即使在使用5%的有偏数据进行训练的情况下也能够输出预测,并且比最先进的模型要准确得多。通过众包,我们验证了语境向量的中立性,并表明语境过滤会导致意识形态集中,从而允许预测分布外的例子。
We propose a novel supervised learning approach for political ideology prediction (PIP) that is capable of predicting out-of-distribution inputs. This problem is motivated by the fact that manual data-labeling is expensive, while self-reported labels are often scarce and exhibit significant selection bias. We propose a novel statistical model that decomposes the document embeddings into a linear superposition of two vectors; a latent neutral context vector independent of ideology, and a latent position vector aligned with ideology. We train an end-to-end model that has intermediate contextual and positional vectors as outputs. At deployment time, our model predicts labels for input documents by exclusively leveraging the predicted positional vectors. On two benchmark datasets we show that our model is capable of outputting predictions even when trained with as little as 5% biased data, and is significantly more accurate than the state-of-the-art. Through crowd-sourcing we validate the neutrality of contextual vectors, and show that context filtering results in ideological concentration, allowing for prediction on out-of-distribution examples.