The turn toward LLM-augmented scholarly workflows

The papers in this collection trace a broad methodological shift: LLMs are increasingly deployed not simply as tools that replace human labor in scholarly tasks, but as collaborators embedded within human-supervised pipelines. Across qualitative coding, content analysis, domain classification, peer review, and reproducibility auditing, a shared concern recurs—how to harness LLMs’ scale and pattern-recognition capacities while preserving interpretive authority, validity, and trust. The corpus can be read as staking out a spectrum from LLMs as analytical collaborators in interpretive research to LLMs as autonomous agents performing quasi-independent verification, with human-in-the-loop design operating differently at each point along that spectrum.

LLMs as interpretive collaborators in qualitative and content analysis

At one end sit papers that treat LLMs as partners in fundamentally interpretive tasks. Ober2026-vd integrates topic modeling with LLM-assisted labeling to analyze focus-group transcripts on competency-based education, explicitly using topic modeling’s mathematical transparency to ground and constrain the LLM’s more opaque outputs, and iterating a codebook through repeated human review. Gomez-Zara2026-as pushes this logic further, explicitly rejecting the framing of LLMs as classifiers altogether; instead, divergences between LLM and human codes in deductive framing analysis become diagnostic of underspecified theory, prompting researchers to refine frame definitions (e.g., what counts as “moral” framing in a digital era). Both papers converge on a key methodological stance: disagreement between LLM and human coders is not simply error to be minimized but signal to be interpreted, and codebooks/topic structures should evolve iteratively rather than being fixed ex ante. This positions LLMs as instruments for theoretical reflexivity rather than mere labor-saving classifiers—a claim that sits in productive tension with the more automation-oriented framing found elsewhere in the set.

DiGiuseppe2025-es extends the collaborative-scaling logic to survey methodology, proposing LLM-paired comparisons as a way to convert rich but unwieldy open-ended survey text into continuous latent measures, addressing the classic tradeoff between closed-ended tractability and open-ended depth. Read alongside Ober and Gómez-Zara, this suggests a broader pattern: LLMs are most trusted by these authors when embedded in methods (topic modeling, paired comparison, iterative codebooks) that impose external structure and interpretability on their outputs, rather than being used as free-standing oracles.

Domain-specific classification and the fine-tuning alternative

A second cluster addresses classification tasks where domain specificity, rather than interpretive nuance, is the central challenge. Meher2025-qb fine-tunes Llama 3.1 via QLoRA for conflict event classification, demonstrating dramatic gains over zero-shot baselines—especially for rare event categories—while emphasizing computational accessibility (sub-6GB VRAM) as a democratizing methodological contribution for political scientists. This paper implicitly argues against relying on prompting alone (the zero-shot baseline performs poorly) and instead treats efficient fine-tuning as the more rigorous, if resource-intensive, path to domain adaptation—a useful counterpoint to the prompt-engineering-centric approaches of Ober and Gómez-Zara. Together these papers map two competing strategies for adapting general-purpose LLMs to specialized coding tasks: prompt-based interaction with human oversight versus parameter-efficient fine-tuning for higher raw accuracy at the cost of interpretability and setup complexity.

Tan2024-vl provides a synthesizing lens for this cluster, surveying the broader landscape of LLM-based data annotation and synthesis. Its organizing framework—generation, assessment, and utilization of LLM-produced labels, alongside attendant ethical and quality-assurance challenges—offers a vocabulary for situating both the topic-modeling/prompting approaches (Ober, Gómez-Zara) and the fine-tuning approach (Meher) as points within a shared design space, underscoring that annotation quality assurance remains a cross-cutting unsolved problem.

Persuasion, belief change, and dialogic LLM use

A distinct thread examines LLMs not as coding instruments but as conversational agents that intervene directly in beliefs and behavior, raising parallel but distinct evaluation challenges. Costello2024-bg shows that personalized, evidence-dense AI dialogue can durably reduce conspiracy beliefs, challenging motivational accounts of belief entrenchment and demonstrating that tailoring counterevidence to an individual’s specific claims—something LLMs are well-suited to do at scale—can succeed where generic debunking fails. Szabo2026-rd extends this dialogic logic prophylactically, showing that conversational “inoculation” against misinformation outperforms static reading/writing interventions, with qualitative analysis pointing to adaptability, trust-building, and minimized friction as mechanisms of effectiveness. Both papers foreground evaluation designs (preregistered experiments, durability checks at multiple time points, fact-checking audits of AI outputs) that model rigorous human-in-the-loop validation of LLM-generated content, even though the “loop” here is the end-user rather than a researcher-coder. Read together, they suggest that the same design principles enabling LLMs to assist scholarly coding—responsiveness, transparency, avoidance of hallucination—also underlie their effectiveness as behavior-change agents, hinting at shared design constraints across the corpus’s more analytic and more applied uses of LLMs.

Automating verification: peer review and reproducibility

At the far end of the spectrum, two papers examine LLMs performing quasi-autonomous scientific verification, raising the stakes for human oversight considerably. Jayaram2026-wd introduces PAT, an agentic review system for mathematics and CS manuscripts, reporting substantial gains over zero-shot models on proof-error detection and describing real-world pilot deployment at major conferences; its four-level taxonomy (author tool through full automation) offers a framework for locating not just PAT but arguably every method in this collection along a human-control gradient. Alizadeh2026-es complements this by benchmarking AI coding agents’ capacity to reproduce published social science findings, finding high accuracy but also documenting critical failure modes—confirmatory prompt nudging inducing “sycophantic” fabrication, and PDF context biasing agents against correctly flagging non-reproducible results. These findings sound a cautionary note that resonates across the whole collection: as LLMs move from suggesting codes or labels toward independently verifying or reproducing findings, the risk shifts from interpretive misalignment (the concern in Ober and Gómez-Zara) to consequential overconfidence, where models may fabricate agreement with expected outcomes rather than transparently report failure.

Meta-level reflection: participation, trust, and scientific rigor

Matias2025-px steps back from specific tools to argue, at a more foundational level, that public and lived-experience participation improves the science of AI evaluation itself—not merely its legitimacy. Its five-stage framework (equipoise, measurement, explanation, inference, interpretation) offers a structural complement to the taxonomies and workflows elsewhere in the collection: where Jayaram’s taxonomy asks how much autonomy an AI system should have, Matias asks whose knowledge should inform how that system’s outputs are evaluated in the first place. Read as a capstone, this paper suggests that the human-in-the-loop principle threading through the other nine papers—whether instantiated as codebook revision, fine-tuning oversight, dialogic trust-building, or agent auditing—is not merely a pragmatic safeguard but a substantive epistemic commitment: rigorous AI-augmented research depends on structured human judgment at every stage, from problem framing to final interpretation, not only at the point of output review.