From Procedure to Practice

The starting move shared across this cluster is a refusal to treat AI safety evaluation as a purely technical exercise. Gillespie2026-aa makes this argument most directly for red-teaming: adversarial probing of generative AI is not a neutral instrument but a sociotechnical system, shot through with value judgments about what counts as harm, organized through specific labor arrangements, and productive of real psychological costs for the people who do it. Matias2025-px generalizes the same insight to AI evaluation writ large, arguing that because AI systems are sociotechnical artifacts—embedded in organizational, political, and social layers—their reliability cannot be assessed by technical measurement alone. Read together, these two papers establish the topic’s foundational claim: safety evaluation, whether adversarial testing or scientific measurement, is a social practice whose rigor depends on who performs it, under what conditions, and with what forms of accountability.

Labor, Harm, and the Hidden Workforce

Gillespie2026-aa extends the content-moderation literature (Gillespie, Roberts, Gray and Suri) into the red-teaming context, showing how the same patterns—outsourcing, precarity, opaque value-setting by internal teams unrepresentative of affected populations, and inadequate mental-health support—are being reproduced with little public scrutiny. Its account of secondary trauma and moral injury among red-teamers, and its skepticism toward claims that this labor will simply be automated away, foreground a dimension almost entirely absent from technical safety discourse: the embodied, often invisible human work that underwrites claims of model safety. This labor-centered lens sits in productive tension with Jayaram2026-wd, which describes an agentic system (PAT) explicitly designed to automate a cognate evaluative labor—scientific peer review—at scale, reporting substantial gains in error detection and high author satisfaction across thousands of submissions. Where Gillespie2026-aa warns that automation narratives distract from the urgent well-being and governance needs of current evaluators, Jayaram2026-wd offers a live case of evaluation work being restructured by AI orchestration, complete with its own emergent taxonomy of human-AI role division (author tool, reviewer tool, supporting reviewer, total automation). The juxtaposition raises a question the topic as a whole leaves open: whether such automation genuinely relieves overburdened human evaluators or simply displaces the same sociotechnical concerns—whose values are embedded in the reviewing agent, who bears responsibility for its errors, what happens to the reviewers it supplements—into a new, less visible layer.

Public Involvement as Scientific Improvement

Matias2025-px pushes the sociotechnical framing toward a constructive program: participatory science methods can make evaluation more rigorous, not merely more legitimate. Its five-stage model—equipoise, measurement, explanation, inference, interpretation—offers a structured account of where lived-experience expertise catches what professional evaluators miss, illustrated by cases like the HRDAG/ACLU reanalysis of the Allegheny Family Screening Tool and community relabeling of police complaint data. This directly complements Gillespie2026-aa’s critique of internally defined, proprietary harm taxonomies in red-teaming: both papers converge on the idea that concentrating value judgments among unrepresentative insiders (whether corporate red-teaming staff or professional evaluators) degrades the epistemic quality of safety claims, not just their perceived fairness. Where Gillespie2026-aa diagnoses the problem and calls for an empirical research agenda, Matias2025-px supplies methodological scaffolding for how public participation could be operationalized without sacrificing scientific standards of generality and reliability.

Transparency as a Sociotechnical Safety Practice

Rauchfleisch2026-fa extends the topic’s public-facing dimension in a different but related direction: rather than involving the public in evaluation design, it treats disclosure to the public during interaction as itself a safety intervention, and subjects that intervention to the same empirical scrutiny Matias2025-px calls for more broadly. Its finding that identity labels (the EU AI Act’s preferred remedy) do little to blunt persuasion, while disclosing persuasive intent roughly halves it, exemplifies the broader argument that meaningful safety practice must attend to the social and psychological mechanisms underlying harm rather than settling for procedural compliance. This resonates with Gillespie2026-aa’s insistence that red-teaming’s current form—compliance-oriented, opaque, market-driven—risks becoming entrenched before its actual mechanisms and effects are understood. Both papers warn against safety practices that satisfy a regulatory or reputational checklist (a disclosure label, a red-teaming report) without empirical evidence that they alter the underlying sociotechnical dynamics they are meant to address.

The Arc: Toward an Empirical, Participatory Science of Safety

Across the four papers, a coherent trajectory emerges. Gillespie2026-aa and Matias2025-px together establish that AI safety evaluation—red-teaming and broader assessment alike—is inescapably social, demanding attention to labor, values, and participation that purely technical framings erase. Rauchfleisch2026-fa demonstrates, through rigorous field experimentation, what it looks like to test a safety practice’s actual sociotechnical effects rather than assume its adequacy from design intent alone—an instantiation of the “science of AI evaluation” Matias2025-px calls for. Jayaram2026-wd, meanwhile, stands as a limit case: an ambitious, high-throughput automation of evaluative labor that has not yet been subjected to the sociotechnical scrutiny—regarding whose judgments are embedded in the reviewing agents, what happens to human reviewers’ expertise and workload, and what harms might go unexamined—that the other three papers insist must accompany any evaluation practice claiming legitimacy. Together, the cluster argues that the future credibility of AI safety work depends less on more sophisticated technical procedures than on making visible, and answerable to the public, the human labor, values, and participatory processes that any evaluation—automated or not—ultimately rests upon.