Data Quality in web-based HIV/AIDS research: Handling Invalid and Suspicious Data.

Data Quality in web-based HIV/AIDS research: Handling Invalid and Suspicious Data.
复制标题

DOI:
10.1177/1525822x12443097
复制
发表时间:
2012-08-01
期刊:
影响因子:
1.7
通讯作者:
Strecher VJ
Strecher VJ
中科院分区:
法学4区
文献类型:
--
作者:
Bauermeister J;Pingel E;Zimmerman M;Couper M;Carballo-Diéguez A;Strecher VJ

文献摘要

被引文献

相似文献

无效数据可能会影响数据质量。我们研究了如何决定处理这些数据可能会影响互联网使用和艾滋病毒的风险行为之间的关系,在一个样本的年轻男性与男性发生性关系(YMSM)。在三个月的时间里,我们记录了548个条目,并创建了6个分析组(即,全样本、最初标记为有效的条目、可疑条目、误标记为可疑的有效案例、欺诈性数据和总有效案例)。我们比较了这些群体的样本的组成和他们的二元关系。41例被标记为无效,影响我们估计的统计精度,但不影响变量之间的关系。另外62个案例被标记为可疑条目,并发现有助于样本的多样性和观察到的关系。使用我们的最终分析样本(N = 447; M = 21.48岁,SD = 1.98),我们发现关于数据排除的非常保守的标准可能会阻止研究人员观察到真正的关联。我们讨论的影响,数据质量的决定和它的影响,未来的艾滋病毒/艾滋病网络调查的设计。
Invalid data may compromise data quality. We examined how decisions taken to handle these data may affect the relationship between Internet use and HIV risk behaviors in a sample of young men who have sex with men (YMSM). We recorded 548 entries during the three-month period, and created 6 analytic groups (i.e., full sample, entries initially tagged as valid, suspicious entries, valid cases mislabeled as suspicious, fraudulent data, and total valid cases) using data quality decisions. We compared these groups on the sample’s composition and their bivariate relationships. Forty-one cases were marked as invalid, affecting the statistical precision of our estimates but not the relationships between variables. Sixty-two additional cases were flagged as suspicious entries and found to contribute to the sample’s diversity and observed relationships. Using our final analytic sample (N = 447; M = 21.48 years old, SD = 1.98), we found that very conservative criteria regarding data exclusion may prevent researchers from observing true associations. We discuss the implications of data quality decisions and its implications for the design of future HIV/AIDS web-surveys.