Stagnaro, M. N., Druckman, J., Berinsky, A. J., Arechar, A. A., Willer, R., & Rand, D. G. (2025). Representativeness and response validity across nine opt-in online samples. PsyArXiv. https://doi.org/10.31234/osf.io/h9j2d_v2

View paper

Summary

This paper delivers a comprehensive empirical comparison of nine widely used opt-in, non-probability online survey samples (total N=13,053), collected between 2022 and 2023 using identical instruments on each platform’s default settings. The authors evaluate samples along three axes — response validity (attentiveness, effort, honesty, non-speeding, non-attrition), representativeness (against the 2020 ANES benchmark, including attitudes and experimental treatment effects), and professionalism (survey frequency, device type). Their central argument is that opt-in samples must not be treated as an interchangeable category: they vary dramatically, and there is a systematic trade-off in which the most representative samples (those employing demographic quotas) tend to show the lowest response validity, plausibly driven by heavy mobile-phone usage. They advance a practical fix — two trivial front-end attention checks — and offer decision-oriented guidance keyed to research purpose.

Key Contributions

  • Most comprehensive side-by-side comparison to date of nine opt-in online samples using identical instruments across validity, representativeness, and professionalism.
  • Documents a systematic representativeness–validity trade-off and identifies mobile-phone usage as a likely mechanism.
  • Shows that two low-cost front-end attention checks improve validity without harming representativeness.
  • Demonstrates that post-stratification weights recover main treatment effects but not heterogeneous treatment effects in non-representative samples.
  • Provides concrete guidelines for platform selection based on study goals (attention, social/political content, heterogeneous effects, cost).
  • Provides evidence that Open Mechanical Turk is dominated on every dimension and should be avoided.

Methods

The authors recruited 13,053 participants across nine platforms (CR Toolkit, CloudResearch Connect, Connect NR, Forthright, Lucid, Open M-Turk, Prolific, Prolific NR, Stanford CRSTAL), each administered an identical survey containing demographics, attitudes, seven attention checks, an effort task, an honesty self-report, speeding/attrition measures, platform-experience questions, and two canonical experiments (welfare/poor framing; Asian-disease risky-choice framing). An aggregate response-validity index averaged the five validity dimensions. Representativeness was measured as summed absolute deviations from the 2020 ANES probability sample, disaggregated by partisanship. Heterogeneous treatment effects were tested by party, race, age, and gender. The team further examined progressive attention-check filtering, post-stratification weighting, and within-platform contrasts (Prolific vs. Prolific NR; Connect vs. Connect NR) to separate platform from sampling effects.

Findings

  • Samples cluster into three validity tiers: high (Connect, Connect NR, Prolific, Prolific NR, CRSTAL), middle (Forthright, CR Toolkit), and low (Lucid, Open M-Turk).
  • Two trivial front-end attention checks raised mean validity from .771 to .822, most improving the lowest-validity samples.
  • Quota-heavy samples (Lucid, Forthright) best matched ANES demographics and attitudes, though all samples deviated more for Republicans and notably on gun control.
  • Filtering on up to two attention checks did not meaningfully harm representativeness (Lucid most sensitive at higher filter levels).
  • The welfare/poor framing main effect replicated everywhere, but heterogeneous effects (by party, race, age) surfaced only in the more representative samples.
  • The risky-choice framing effect replicated with no clear demographic moderation, suggesting cognitive experiments are less sample-dependent.
  • Weights restored main effects (even in Open M-Turk) but never recovered heterogeneous treatment effects.
  • Professional respondents concentrated in less representative samples and correlated with higher validity; representative samples relied heavily on mobile respondents (~70% on Lucid/Forthright vs. <10% elsewhere).
  • Expected identity correlations (Black respondent ideology/party; male bias in Trump support) appeared only in representative samples, raising concerns about identity misreporting elsewhere.

Connections

This paper is a foundational methodological benchmark for the broader survey-methodology-validity literature. It connects most directly to work probing data quality and inattention in online panels, such as DiGiuseppe2025-es and Luhring2025-od, and to studies concerned with the validity of survey-based measurement and inference more generally, such as Ulloa2024-jm.

Podcast

A research-radio episode discusses this paper: 🎧 MP3 · Spotify · Apple Podcasts