Gillespie, T., Shaw, R., Gray, M. L., & Suh, J. (2026). AI red-teaming is a sociotechnical problem. Communications of the ACM. https://doi.org/10.1145/3731657

View paper

Summary

This conceptual essay argues that AI red-teaming — the adversarial probing of generative AI systems for vulnerabilities, harmful outputs, and biases — should be understood as a sociotechnical practice rather than a narrowly technical procedure. The authors contend that red-teaming has been rapidly normalized as a central AI safety and policy mechanism even though the public knows little about how it is conducted, by whom, or according to whose values. Drawing an extended parallel with the history of commercial content moderation, they surface three underexamined dimensions: the value judgments embedded in defining harms, the labor arrangements that organize the work, and the psychological toll it exacts on workers. Their central move is a warning and a call: before red-teaming’s current opaque, market-driven form becomes entrenched, a coordinated interdisciplinary research network should study it empirically.

Key Contributions

  • Reframes AI red-teaming as a sociotechnical problem rather than a purely technical or compliance issue.
  • Develops a structured comparison between red-teaming and commercial content moderation, extracting transferable lessons about values, labor, and worker harm.
  • Names psychological risks specific to red-teaming, including moral injury arising from sustained adversarial roleplay and transgressive imagination.
  • Issues a programmatic call for a cross-disciplinary research network (computer science, social science, humanities, law).
  • Offers a critical vocabulary — value judgments, labor politics, well-being — to organize future empirical and policy work on AI safety labor.

Methods

A conceptual and critical essay, explicitly not based on internal Microsoft information, synthesizing the authors’ prior work on Responsible AI labor, generative AI politics, and participatory governance. The argument proceeds by comparative analysis (red-teaming vs. content moderation) and engagement with science and technology studies, labor studies, psychology, and design. It reviews public-facing materials from major AI companies (OpenAI, Anthropic, Google, Microsoft), policy documents (Executive Order 14110), and events such as DEFCON 2023’s Generative Red Team.

Findings

  • Definitions of red-teaming remain fuzzy, overlapping with evaluation, bug bounties, penetration testing, and “ethical hacking,” and are institutionally unsettled.
  • Internal red-teamers often lack the sociocultural, linguistic, and ethical expertise to identify diverse harms, and may be constrained by NDAs and corporate incentives — leaving value judgments with designers who do not reflect affected users.
  • Red-teaming labor is increasingly outsourced to vendors and crowdworkers, replicating the labor arbitrage, weak protections, and precarity of content moderation.
  • Volunteer and event-based red-teaming (e.g., DEFCON) can broaden participation but risks extractive reliance on marginalized communities and does not scale.
  • Workers face secondary traumatic stress, PTSD-like symptoms, and moral injury from repeated exposure to harmful content and from inhabiting adversarial personas.
  • Existing well-being supports (EAPs, content warnings, opt-outs) are inconsistently applied and undermined by NDAs, performance pressure, and surveillance dynamics.
  • Claims that red-teaming will be automated away are misleading and obscure the human labor involved rather than eliminating it.

Connections

This essay’s participatory and event-based angle on red-teaming connects to work on broadening evaluation beyond corporate insiders, such as Nguyen2026-vm and Ng2026-og. Its attention to the hidden labor and well-being costs of AI work — extending the “ghost work” tradition into generative AI safety — resonates with research on the labor impacts of AI, though the specific papers grouped under that topic here appear to address public perceptions and downstream effects rather than the frontline red-teaming labor this essay foregrounds.

Podcast

A research-radio episode discusses this paper: 🎧 MP3 · Spotify · Apple Podcasts