The methodological turn: LLMs as instruments of content analysis
The organizing anchor for this cluster is Marino2024-2fbc690f, which names the object under study directly: the “LLMs-in-the-loop pipeline,” a workflow in which generative models are chained into fine-tuning, embedding, and labeling stages and must therefore be validated at each stage rather than as a monolith. This essay crystallizes a problem that recurs across nearly every other paper filed here — that LLMs are simultaneously classifiers, feature extractors, and interpreters, and that the validation logic developed for single-purpose supervised models (train/test splits, inter-coder reliability) does not transfer cleanly to systems that are general-purpose, prompt-sensitive, and subject to silent vendor-side change. Balluff2026-if radicalizes this concern into an outright critique of “unreflective” LLM adoption in political communication research, arguing that convenience has outrun methodological rigor and that researchers routinely ignore reproducibility, environmental cost, and corporate dependency in their haste to swap in ever-larger commercial models. Tan2024-vl provides the disciplinary counterpart from computer science, surveying the rapidly growing literature on LLM-based annotation and synthesis and confirming that the field still lacks settled standards for assessing when a generated label or synthetic datapoint is trustworthy. Read together, these three pieces frame the central methodological question this whole topic circles: not whether LLMs can annotate, cluster, or classify communication data — they clearly can — but under what conditions their outputs are valid stand-ins for human judgment.
Classification and annotation: performance, bias, and validity
A cluster of papers puts LLM-based classification to direct empirical test against human-coded or supervised baselines. Bailard2024-pj fine-tunes DeBERTa to classify collective-action frames in over 500,000 Proud Boys Telegram messages, demonstrating how transformer-based classification can be married to time-series causal inference (Granger causality, impulse-response) to link online discourse to offline violence — a template for LLM-assisted classification feeding downstream social-scientific modeling. Meher2025-qb pushes efficient fine-tuning further, showing that QLoRA-adapted Llama 3.1 can outperform BERT-style encoders like ConfliBERT on conflict-event classification while running on consumer hardware, reframing LLM adoption as a question of accessibility rather than raw capability. Larsson2026-ro and Marino2026-he both use GPT-4o zero-shot classification (validated against human coders) to scale sentiment and entity-engagement coding across a decade of Norwegian and three years of Brazilian Facebook data respectively, illustrating that LLM-assisted classification is now a standard tool for longitudinal, non-English political communication research — with Marino2026 explicitly building on the pipeline logic of Marino2024.
Two papers interrogate whether LLM annotation reproduces or distorts human judgment. Brown2025-jk finds that demographic disagreement between LLMs and human annotators on contentious labeling tasks (toxicity, offensiveness, politeness) is dataset-specific rather than model-specific, and that item difficulty (label entropy) swamps demographic bias as a predictor of agreement — a genuinely reassuring, if narrow, validity result. DeVerna2025-dl complicates any general optimism, showing that even reasoning- and search-augmented LLMs fail at political fact-checking absent curated retrieval context, and that search-enabled models exhibit a citation skew toward left-leaning sources — a finding with direct implications for anyone using “LLM as fact-checker” in a content-analysis pipeline. Paci2025-ag extends validity concerns from surface classification to pragmatics, showing current LLMs struggle to correctly interpret implicatures and presuppositions in Italian political discourse, exposing a ceiling on how far prompt-based coding can substitute for expert qualitative interpretation of implicit meaning. Alizadeh2026-es takes validation furthest upstream, benchmarking whether AI coding agents (Claude Code, Codex) can reproduce published social-science findings outright — with strong headline accuracy but revealing new failure modes (sycophantic specification search under confirmatory prompting) that are directly relevant to LLM-in-the-loop validation more broadly.
Embeddings, clustering, and the question of what similarity means
A second strand treats embeddings, not classification labels, as the primary object of methodological scrutiny. Giglietto2024-cbeb3f70 directly compares OpenAI’s text-embedding-3-large against the Italian-specific UmBERTo model for clustering political news, finding the general-purpose LLM embedding consistently superior and validating cluster quality via a GPT-4o-mini-based adaptation of Grimmer and King’s coherence metric — a result that feeds directly into the clustering stage of the Marino2024 pipeline. Fan2025-ut tackles a more fundamental confound: pretrained embeddings encode spurious source/language signal that biases downstream similarity measures, and the paper shows that linear concept erasure (LEACE) can remove this confounding cheaply and reliably, a finding with obvious relevance to any cross-source or cross-lingual clustering study in this cluster. Le-Mens2025-qz proposes an entirely prompt-based alternative to embedding-based scaling, asking LLMs directly to place texts on ideological dimensions and averaging across many queries — a “measurement by asking” approach that sidesteps clustering altogether. Lai2024-to and Bruns2025-fz extend the embedding logic to new representational objects: video-level ideology estimated from Reddit sharing behavior, and “practice mapping” of multimodal network interactions as an alternative to conventional network visualization’s “hairball” problem.
Multimodal extensions: images, video, and mixed media
The move from text to multimodal content is a distinct sub-front. Achmann-Denkler2026-lx and Arminio2025-tw both show multimodal LLMs (GPT-4o, VLLMs) outperforming specialized computer-vision pipelines — for face recognition/person-counting in campaign imagery, and for connotative (rather than merely denotative) semantic clustering of climate-change images on Instagram — with Arminio2025 explicitly foregrounding Barthesian connotation as the theoretical target that CNN-based pipelines cannot reach. Arora2025-tx generalizes this logic to framing analysis itself, arguing that automated frame extraction has wrongly confined itself to text and fixed frame sets, and proposing multi-modal, discourse-analytic frame extraction across text, image, and their interaction.
Beyond classification: narrative, alignment, and latent constructs
A distinctive contribution of this literature is using LLMs not just to replicate existing coding schemes but to operationalize theoretical constructs that were previously too labor-intensive to measure at scale. Elfes2026-jb operationalizes Greimas’ Actantial Model via an LLM to measure “narrative polarisation,” finding that videos are far more polarized than the comments beneath them — a genuinely novel measurement enabled by LLM annotation. Waight2025-al builds a comparably ambitious pipeline to measure cross-lingual “narrative similarity,” explicitly arguing that narrative overlap is a distinct estimand from lexical, topical, or semantic similarity, and validating an LLM-based claim-distillation method against exact-reuse, topic-modeling, and Relatio baselines. Sarmiento2025-as pursues a related but more inductive goal, combining ML, network analysis, and NLP to surface emergent frames in polarized discourse without predefined categories. Ober2026-vd shows how topic modeling and LLM prompt engineering can be integrated into a human-in-the-loop qualitative workflow for interview transcripts, explicitly defending topic modeling’s mathematical transparency as a complement to LLM opacity. Fan2026-af widens the aperture further still, reviewing six computational approaches (including language-based/embedding models) for analyzing the temporal structure of digital trace data, positioning neural embeddings as a promising but data-hungry tool for capturing sequence-level behavioral patterns. Lee2026-je demonstrates a related but more unsettling capability: LLMs can infer users’ political alignment even from nonpolitical text, exploiting subtle lexical and cultural cues, raising the stakes on what “content analysis” can now extract from ostensibly innocuous discourse.
LLMs as objects and agents of communication, not just tools
Several papers use these same methods to study LLMs themselves as communicative actors, which both showcases and stress-tests the toolkit. Hackenburg2025-dj, DiGiuseppe2026-pu, and Lin2025-xp form a tight empirical sequence on conversational AI persuasion — establishing that information density and post-training (not scale or personalization) drive persuasive effects, that perceived political bias attenuates persuasion, and that AI dialogue can shift voter preferences across four national contexts while exhibiting a partisan accuracy asymmetry. Orlando2025-ul and Allen2025-ot extend the same agentic lens to platform-level manipulation and independent auditing: generative agent-based modeling shows LLM agents can spontaneously reproduce coordinated inauthentic behavior, while Allen and Tucker’s commentary champions LLM-powered browser-extension experiments as a “platform-independent” alternative to increasingly closed APIs. Waight2026-ts and Triedman2025-uy both audit LLM outputs and AI-generated content for embedded political bias — state media influence laundered into model responses, and Grokipedia’s derivative, differentially unreliable sourcing relative to Wikipedia — using embedding similarity and citation analysis as their evidentiary base. Nguyen2026-vm closes the loop by studying how news media frame LLMs themselves, combining computational and manual content analysis of 18,000+ articles to map anthropomorphic framing of the very technology this whole cluster deploys as method.
Reflexivity, epistemology, and infrastructure
A final set of papers steps back to interrogate the conceptual and institutional conditions of this methodological moment rather than deploying LLMs directly. Dierickx2026-tw asks what counts as a “fact” when content is generated rather than retrieved, proposing “emergent facts” as a new epistemic category — a direct theoretical complement to the empirical fact-checking failures documented in DeVerna2025. Sadler2025-vu argues, in a similar reflexive register, that disinformation research’s loose use of “narrative” needs hermeneutic and narratological grounding, offering a conceptual counterpart to the more computational narrative-measurement projects of Elfes2026 and Waight2025. Manovich2026-ih takes the furthest remove, theorizing generative AI as an artistic medium with built-in “media cognition” — a reminder that the same models powering this cluster’s classification pipelines carry structural biases (aesthetic conservatism, style-content entanglement) that are relevant wherever LLMs are asked to interpret or generate meaning rather than merely label it. Finally, Peters2026-mo and Bruns2026-pn address the infrastructural precondition for all of this work: data access. Peters shows how “data quality” became a politically contested category in the EU’s Digital Services Act negotiations, while Bruns documents how “clean room” access regimes (replacing CrowdTangle) structurally privilege quantitative, code-based, English-language research — exactly the kind of LLM-in-the-loop pipelines this topic catalogs — while excluding qualitative and Global-South scholarship. Nenno2025-xa and Anwar2024-34dba628 round out the empirical base, using computational news-values detection and a systematic review of Facebook Reactions research to show how even “simple” platform signals require the same validation scrutiny as LLM outputs when generalized across cultural and linguistic contexts.
Across these forty papers, the arc is consistent: enthusiasm for LLMs’ classification and clustering power, followed swiftly by demonstrations that this power is uneven (strong on explicit sentiment and lexical politics, weak on pragmatics and fact-checking without curated context), and capped by a maturing methodological literature — anchored in Marino2024 and Balluff2026 — insisting that every stage of an LLM-in-the-loop pipeline demands its own validation, its own reflexivity about infrastructural constraints, and its own reckoning with what, precisely, the model is being asked to measure.