Reframing Red-Teaming and Evaluation as Sociotechnical Work
The core move shared across this cluster is a refusal to let “safety evaluation” remain a purely technical category. Gillespie2026-aa and Unknown2025-qj both target red-teaming specifically, arguing that adversarial probing of generative AI is inseparable from questions of who is doing the probing, under what institutional arrangement, and whose values determine what counts as a harm. Gillespie2026-aa makes the case historically, tracing red-teaming’s rapid, opaque institutionalization back to the earlier normalization of commercial content moderation—the same pattern of outsourced labor, internally defined harm categories, and invisible psychological cost. Unknown2025-qj arrives at a similar diagnosis from ethnographic fieldwork (26 interviews, participant observation at three public red-teaming events), showing empirically how organizational context and framing shape what vulnerabilities even become visible, and gesturing toward red-teaming’s roots in cybersecurity and military practice as a lineage still shaping its present form. Together these two papers establish the topic’s foundational claim: red-teaming is not a neutral technical procedure but a value-laden, labor-intensive, institutionally contingent practice that has been allowed to consolidate without public scrutiny.
Matias2025-px extends the same sociotechnical premise from red-teaming to evaluation more broadly. Where Gillespie2026-aa worries about who is harmed by red-teaming labor, Matias2025-px asks who is missing from the evaluation process and argues—crucially—that this is not just an ethical or political deficiency but a scientific one. The paper’s five-stage framework (equipoise, measurement, explanation, inference, interpretation) gives structural specificity to a claim Gillespie2026-aa and Unknown2025-qj make more diagnostically: that excluding lived-experience expertise doesn’t just raise legitimacy concerns, it produces worse science, as shown in the Allegheny Family Screening Tool and Chicago police-complaint reanalyses where community-driven relabeling surfaced biases and hidden harms invisible to standard technical metrics.
Labor, Harm, and the Limits of Volunteerism
A throughline connecting Gillespie2026-aa and Unknown2025-qj is skepticism toward participation as an unproblematic good. Gillespie2026-aa is explicit that public and volunteer red-teaming events (e.g., DEFCON-style exercises) can broaden whose harms get surfaced, but risk becoming extractive—relying on marginalized communities’ labor without corresponding protection, compensation, or well-being infrastructure, echoing the ghost-work dynamics of data labeling and content moderation. This tempers the more optimistic register of Matias2025-px, which emphasizes the epistemic gains from lived-experience participation but says comparatively little about the working conditions or psychological toll such participation might impose on contributors. Read together, the two papers suggest an unresolved tension in this literature: participatory and public-facing safety work is simultaneously the corrective to narrow technical evaluation and a potential new site of precarious, under-protected labor if institutionalized carelessly—a tension Unknown2025-qj’s field data on public red-teaming events is well positioned to probe further but does not fully resolve.
The Structural Threat: Who Controls the Evidence Base
Bak-Coleman2025-pm operates one level up from the practice of evaluation itself, addressing the political economy that determines whether any independent sociotechnical inquiry into AI and platforms is possible at all. Its argument—that tech companies control the data, funding, and access needed to study their own products, mirroring tobacco- and fossil-fuel-style evidentiary capture—supplies the structural backdrop against which Gillespie2026-aa’s complaint about internally defined, proprietary harm categories and Matias2025-px’s case for public involvement both become more urgent. If companies can suppress internal research, restrict data access, and stage “performative” collaborations, then the participatory correctives Matias2025-px proposes and the public-interest reframing Unknown2025-qj documents are perpetually at risk of being absorbed or neutralized by the same industry actors whose products are under scrutiny. Bak-Coleman2025-pm thus supplies the topic’s institutional stakes: without external mechanisms—mandated data access, disclosure requirements, independent funding—calls for sociotechnical, participatory evaluation risk remaining aspirational rather than structurally guaranteed.
A Counterpoint: The Pull Toward Automation
Jayaram2026-wd sits at an angle to the rest of this cluster and is useful precisely as a foil. Where the other four papers argue for more human, public, and participatory involvement in evaluating AI systems, this paper describes an agentic tool (PAT) built to automate scientific peer review at scale, justified by the claim that human review “cannot scale” to match AI-assisted submission volume. The framing is instructive: it treats evaluation labor as a bottleneck to be engineered away through inference-scaled agentic pipelines, rather than as a site of value judgment, labor politics, or public accountability. Set against Gillespie2026-aa’s warning that automation narratives about red-teaming “distract from urgent labor… concerns about the people currently doing this work,” Jayaram2026-wd exemplifies exactly the kind of purely technical, throughput-oriented solution this topic’s other papers are arguing against—even as it operates in the adjacent domain of scientific verification rather than AI safety per se. Its inclusion sharpens the stakes: as evaluation and review infrastructures scale, the sociotechnical case made by Gillespie2026-aa, Unknown2025-qj, Matias2025-px, and Bak-Coleman2025-pm is precisely the argument that must be made against the default assumption that scale problems are best solved by removing humans—and the publics they represent—from the loop.