Identification, Characteristics and Impact of Faked Interviews in Surveys: An Analysis by Means of Genuine Fakes in the Raw Data of SOEP

Identification, Characteristics and Impact of Faked Interviews in Surveys: An Analysis by Means of Genuine Fakes in the Raw Data of SOEP
复制标题

调查中造假访谈的识别、特征和影响:基于SOEP原始数据中真假的分析

DOI:
--
复制
发表时间:
2003
期刊:
Social Science Research Network
影响因子:
--
通讯作者:
Gert G. Wagner
Gert G. Wagner
中科院分区:
--
文献类型:
--
作者:
Joerg–Peter Schraepler;Gert G. Wagner

文献摘要

被引文献

相似文献

据我们所知,在为数不多的分析假面试对调查结果影响的方法论研究中,大多数都是基于项目学生在“实验室环境”中产生的“人工假”。相比之下,面板数据提供了一个独特的机会来识别实际上是由采访者伪造的数据。通过比较两个波的数据,明确的假货很容易识别。然而,在大多数调查中,没有第二波,因为它们具有纯粹的横截面性质。为了寻找一种不需要两波数据的方法,我们测试了一种名为本福德定律的非传统基准,几位会计师使用该定律来发现欺诈行为。我们的初步结果让我们得出结论,本福德定律可能不是检测虚假数据的有效方法,但它可能是采访过程质量控制的新工具德国社会经济小组研究(SOEP)的原始数据提供了丰富的虚假采访来源,因为它是建立在几个子样本之上的。然而,由于面试官知道小组受访者将在一段时间内再次接受面试,聪明的面试官不会伪造小组面试。事实上,在SOEP的原始数据中,该份额仅占所有记录的0.5%。这些假货用于分析未检测到的假货对调查结果的潜在影响。主要结果是,伪造的记录没有影响的平均值和比例。但在非常罕见的特殊情况下,如果无法检测到伪造品,则相关性和回归系数的估计可能存在偏差。应该注意的是,除了前两波样本E中的一些伪造数据外,伪造数据从未在广泛使用的SOEP中传播。在数据公布之前就发现了假货。
To the best of our knowledge, most of the few methodological studies which analyze the impact of faked interviews on survey results are based on “artificial fakes” generated by project students in a “laboratory environment”. In contrast, panel data provide a unique opportunity to identify data which are actually faked by interviewers. By comparing data of two waves, unequivocal fakes are easily identifiable. However, in most surveys there is no second wave because they have a pure cross-sectional nature. In search of a method which does not need two waves of data we test an unconventional benchmark called Benford’s Law, which is used by several accountants to discover frauds. Our preliminary results let us conclude that Benford´s Law might be not an efficient method for detecting faked data, but it might be a new instrument for quality control of the interviewing process The raw data of the German Socio-Economic Panel Study (SOEP) provide a rich source of faked interviews because it is built on several sub-samples. However, because interviewers know that panel respondents will be interviewed again over the course of time, clever interviewers will not fake panel interviews. In fact, in raw data of SOEP the share is about only 0,5 percent of all records. The fakes are used for an analysis of the potential impact of non detected fakes on survey results. The major result is that the faked records has no impact on the mean and the proportions. But in very rare, exceptional cases there may be a bias in estimates of correlations and regression coefficients if fakes would not be detected. One should note that – except for some fakes in the first two waves of sample E – faked data were never disseminated within the widely-used SOEP. The fakes were detected before the data were released.