Ensuring survey research data integrity in the era of internet bots.

Ensuring survey research data integrity in the era of internet bots.
复制标题

DOI:
10.1007/s11135-021-01252-1
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Halkitis PN
Halkitis PN
中科院分区:
社会科学3区
文献类型:
--
作者:
Griffin M;Martino RJ;LoSchiavo C;Comer-Carruthers C;Krause KD;Stults CB;Halkitis PN

文献摘要

被引文献

相似文献

我们利用基于互联网的调查平台,就COVID-19对美国LGBTQ +人群的影响进行了横断面调查。虽然这种数据收集方法快速而廉价,但由于机器人的渗透,收集的数据需要大量的清理。根据这一经验,我们提供了确保数据完整性的建议。招募于2020年5月7日至8日进行,初步样本为1251份。qualics的调查是通过社交媒体和专业协会列表服务器发布的。在注意到数据差异后,研究人员制定了严格的数据清理协议。2020年6月11日至12日,按照原招聘方式进行第二轮招聘。五步数据清理方案从初始数据集中删除了773项(61.8%)调查,在第一波数据收集中产生了478名参与者的样本。该方案导致从第二波为期两天的数据收集中删除了46项(31.9%)调查,导致第二波数据收集的样本为98名参与者。在验证了为期两天的试点过程在筛选机器人方面是有效的之后,调查重新开始进行第三轮数据收集,共收到709份回复,其中确定了额外的514名(72.5%)有效参与者,并导致额外的194名(27.4%)可能的机器人被删除。最终的分析样本由1090名参与者组成。尽管互联网研究是一种有用而有效的研究工具,尤其是在难以接触的人群中,但尽管调查平台有内置的保护措施,但基于互联网的研究很容易受到机器人和恶作剧应答者的攻击。除了耗尽研究资金外,机器人渗透还威胁到数据的完整性,并可能对边缘人群的研究造成不成比例的损害。根据我们的经验,我们建议使用定性问题、重复人口统计问题和奖励性抽奖等策略来减少恶作剧受访者的可能性。可以采取这些保护措施,以确保数据的完整性并促进对弱势群体的研究。
We used an internet-based survey platform to conduct a cross-sectional survey regarding the impact of COVID-19 on the LGBTQ + population in the United States. While this method of data collection was quick and inexpensive, the data collected required extensive cleaning due to the infiltration of bots. Based on this experience, we provide recommendations for ensuring data integrity. Recruitment conducted between May 7 and 8, 2020 resulted in an initial sample of 1251 responses. The Qualtrics survey was disseminated via social media and professional association listservs. After noticing data discrepancies, research staff developed a rigorous data cleaning protocol. A second wave of recruitment was conducted on June 11–12, 2020 using the original recruitment methods. The five-step data cleaning protocol led to the removal of 773 (61.8%) surveys from the initial dataset, resulting in a sample of 478 participants in the first wave of data collection. The protocol led to the removal of 46 (31.9%) surveys from the second two-day wave of data collection, resulting in a sample of 98 participants in the second wave of data collection. After verifying the two-day pilot process was effective at screening for bots, the survey was reopened for a third wave of data collection resulting in a total of 709 responses, which were identified as an additional 514 (72.5%) valid participants and led to the removal of an additional 194 (27.4%) possible bots. The final analytic sample consists of 1090 participants. Although a useful and efficient research tool, especially among hard-to-reach populations, internet-based research is vulnerable to bots and mischievous responders, despite survey platforms’ built-in protections. Beyond the depletion of research funds, bot infiltration threatens data integrity and may disproportionately harm research with marginalized populations. Based on our experience, we recommend the use of strategies such as qualitative questions, duplicate demographic questions, and incentive raffles to reduce likelihood of mischievous respondents. These protections can be undertaken to ensure data integrity and facilitate research on vulnerable populations.