The arc of inquiry

This collection traces a methodological arc that runs from establishing that LLMs can do the work of human coders, through interrogating what that apparent competence actually captures and misses, to deploying LLMs and their embeddings as substantive instruments for measuring political meaning — framing, narrative, ideology, bias — and finally to reflexive stock-taking about the costs and epistemic politics of the whole enterprise.

Building and validating the pipeline

Marino2024-2fbc690f is close to an ur-text here: it documents an “LLMs-in-the-loop” pipeline (fine-tuned political classifier → embedding-based clustering → LLM-generated cluster labels) for Italian Facebook election content, and names the validation problem that recurs throughout the corpus — general-purpose LLMs don’t specify their competence narrowly enough for classical inter-coder-reliability logic to transfer cleanly. Giglietto2024-cbeb3f70 isolates exactly the embedding stage of that pipeline, comparing OpenAI’s text-embedding-3-large against the Italian UmBERTo for clustering the same corpus and finding the proprietary embedding consistently more coherent — giving methodological teeth to Marino2024-2fbc690f’s choice. Larsson2026-ro shows the same GPT-4/human-validation hybrid working at scale in a different register, classifying a decade of Norwegian party Facebook posts for sentiment to support a substantive negative-campaigning argument.

Where these treat LLM annotation as workable-with-caveats, Gomez-Zara2026-as, Paci2025-ag, and Brown2025-jk push harder on what “validation” is doing. Gomez-Zara2026-as reframes divergence between LLM and human coders applying a framing codebook not as error but as diagnostic of where the codebook itself under-specifies theory — the LLM as “analytical collaborator” rather than benchmarked classifier. Paci2025-ag supplies a harder ceiling: on implicit content (implicature, presupposition) in Italian political speech, even the best model falls twenty points short of an estimated human ceiling. Brown2025-jk asks whether LLM-human disagreement tracks annotator demographics or something else, finding that item difficulty (human label entropy) swamps demographic bias as a predictor — a corrective to assuming annotation failures are principally about whose perspective a model encodes. Meher2025-qb extends validation into fine-tuning rather than prompting, showing that parameter-efficient adaptation (QLoRA) on consumer hardware can match or beat prompting for multi-label conflict-event classification, especially on rare categories.

Embeddings and the geometry of political text

A parallel strand treats the embedding space itself as the object of study. Fan2025-ut shows that pooled sentence embeddings carry confounding structure — source, language — that swamps genuine semantic similarity, and that linear concept erasure recovers dramatic clustering and retrieval gains without degrading out-of-domain performance, complementing Giglietto2024-cbeb3f70’s comparative-model finding with a lesson about what must be stripped out before embedding distance can stand for meaning-distance. Arminio2025-tw extends the embedding question to images, arguing that CNN-based clustering captures denotative content but misses connotative, culturally embedded meaning; routing images through a vision-language model into text before embedding improves both connotative coherence and interpretability. Le-Mens2025-qz offers a more parsimonious alternative: simply asking an instruction-tuned LLM to place a text on a policy dimension and averaging responses reproduces established ideological-scaling benchmarks across languages, substituting for both classical scaling and embedding-based measurement. Lai2024-to extends ideology-scaling logic to video, using Reddit cross-posting behavior to place YouTube videos on a latent left-right dimension — a reminder that the objects of political scaling keep expanding, from legislators to users to media outlets to individual videos, alongside newer LLM-native approaches.

From classification to narrative and framing

A third cluster uses LLMs to extract latent structure that classical content analysis struggled to operationalize. Waight2025-al makes the estimand problem explicit: “narrative similarity” is conceptually distinct from lexical overlap, topic similarity, or embedding similarity, and existing estimators systematically miss the diffusion phenomena researchers care about — an empirical vindication of the estimand-estimator concerns raised more abstractly by Marino2024-2fbc690f and Gomez-Zara2026-as. Elfes2026-jb pushes narrative analysis toward structuralist theory, using an LLM to operationalize Greimas’s Actantial Model on Israeli-Palestinian YouTube discourse, finding that surface narrative divergence between partisan audiences collapses in comments even as deeper motifs preserve partisan difference. Arora2025-tx and Sarmiento2025-as pursue complementary escapes from fixed-frame analysis — the former via multi-modal (text-plus-image) framing of gun-violence news, the latter via an unsupervised ML/network/NLP pipeline for emergent frames in polarized discourse. Achmann-Denkler2026-lx belongs to the same “generalist model outperforms bespoke pipeline” logic in the visual register, with GPT-4o beating specialized computer-vision tools at face recognition and person-counting in campaign Instagram content, again turning a validation question into a substantive finding about concentrated visibility.

LLMs as instruments — and vectors — of political signal

A further set turns LLMs and related architectures into probes of political information, often with troubling implications. Lee2026-je shows LLMs can infer political alignment from ostensibly nonpolitical cultural references, outperforming supervised baselines. DeVerna2025-dl finds that neither scale, reasoning, nor web search closes the political fact-checking gap — only curated retrieval does — and that search-enabled models cite left-leaning (if credible) sources disproportionately. Waight2026-ts extends the bias question upstream to training data, showing state-coordinated media is memorized by commercial LLMs and shapes outputs by query language. Hackenburg2025-dj and DiGiuseppe2026-pu turn to persuasion: the former finds post-training and information density are the dominant levers of AI persuasiveness, at the cost of accuracy; the latter shows mere perception of model bias undercuts persuasive power — together suggesting content-analysis pipelines cannot treat LLM output as a neutral instrument. Bouchaud2026-lr, though about a recommender system rather than a generative model, belongs in this company: X’s embedding space inadvertently encodes a near-linear left-right direction independent of demographics — the same linear-representation logic underlying the deliberate correction in Fan2025-ut, here discovered as unintended consequence.

Reflexivity, reproducibility, and the wider ecosystem

The final cluster steps back to ask what LLM adoption does to the research enterprise itself. Balluff2026-if is the most direct critique, cataloguing reproducibility, environmental, and bias costs across text analysis, synthetic data, and simulation, and counseling a trade-off mindset — a position Giglietto2024-cbeb3f70’s cost reporting and Meher2025-qb’s consumer-hardware fine-tuning both implicitly answer. Tan2024-vl supplies the field-level survey backdrop, organizing LLM-based annotation and synthesis into generation, assessment, and utilization. Alizadeh2026-es asks a forward-looking reproducibility question — can AI coding agents reproduce social-science findings? — documenting both strong success rates and a susceptibility to sycophantic specification search that echoes DiGiuseppe2026-pu’s concern about models telling users what they want to hear. Allen2025-ot and Fan2026-af extend reflexivity to infrastructure: the former situates LLM-based browser-extension methods within a menu of platform-independent designs necessitated by closing API access; the latter argues digital-trace research has under-exploited fine-grained temporal structure, proposing sequence- and process-mining as complements to language-model analysis. Goldberg2026-eb takes the widest lens, asking whether LLM-enabled dialogue, bridging, and moderation systems could reshape digital public squares — an explicitly normative bookend to the mostly instrumental work elsewhere. Marino2026-he closes the loop empirically, reusing GPT-4o classification to show that pro-Bolsonaro Facebook communities are far less stable and echo-chamber-like than assumed. Ahuja2025-ku, finally, stands apart in using no LLM at all, but shares the collection’s underlying concern with translating messy normative categories — autonomy violation, in its case — into operationalizable, auditable factors, the same translation problem Gomez-Zara2026-as and Waight2025-al confront in making “frame” or “narrative” legible to a computational pipeline.