The arc from open APIs to regulated access
The through-line of this collection is a historical arc: an early “Wild West” of loosely governed platform data, a mid-2010s “Golden Age” of sanctioned academic access, and a post-2018 collapse into what Freelon2018-ao first named the “post-API age.” Bruns2019-nr frames the Cambridge Analytica fallout as an “APIcalypse” in which platforms weaponized privacy concerns to frustrate independent scrutiny, a reading extended and periodized by Freelon2024-sc and Giglietto2026-855a54cb, both of which trace Facebook/Meta, Twitter/X, TikTok, and Reddit through parallel cycles of “data philanthropy,” partial API openness, and eventual paywalling or shutdown. Bastos2025-ya and Murtfeldt2025-wu eulogize the specific loss of Twitter’s research infrastructure, with the latter quantifying a real, measurable 13% decline in Twitter-based publications in 2024 — empirical confirmation of the qualitative story others tell. Yang2026-tq’s systematic review shows this is not merely a Twitter phenomenon: across social science disciplines, the share of social-media-based empirical research plateaued and declined after 2022, precisely as API restrictions bit. Into this vacuum steps the EU’s Digital Services Act, and Article 40 in particular, as the collection’s central regulatory hinge — discussed programmatically by Ohme2026-nv, de-Vreese2026-zx, Pierri2026-ib, Philipp2026-tl, and Lukito2026-nb. These pieces largely converge on a shared diagnosis: Article 40 represents a genuine paradigm shift from data access as platform philanthropy to data access as legal entitlement, but its implementation is halting, unevenly enforced, and geographically restricted to the EU, producing what Giglietto2026-855a54cb calls a “data abyss” for Global South and under-resourced researchers.
The texture of non-compliance: audits of the new infrastructure
A cluster of empirical audits tests whether DSA-mandated access mechanisms actually deliver on their promise, and the verdict is largely skeptical. Entrena-Serrano2025-gw finds TikTok’s expanded Research API riddled with unexplained inconsistencies; Rieder2025-ju shows YouTube’s Data API search endpoint is “forgetful by design,” systematically losing findability for older videos in ways that undermine DSA-relevant systemic-risk research; Jurg2025-ur extends this critique to YouTube’s content-moderation transparency during EU elections, finding inconsistent publisher-context labeling and vague removal statements. Tonneau2025-bv uses DSA transparency data itself to expose stark linguistic inequities in human moderator allocation, with Global South languages chronically under-resourced relative to English. Bruns2026-pn and Cullen2026-cb turn attention to the form of access rather than its mere existence: both argue that “clean room” environments (epitomized by Meta Content Library, successor to CrowdTangle) and API designs like CrowdTangle’s constrain not just what data is available but what kinds of researchers, methods, and questions are even possible — privileging quantitative, code-fluent, well-resourced teams over qualitative, comparative, or Majority World scholarship. Peters2026-mo adds a further wrinkle, showing that “data quality” — ostensibly a technical matter — is itself a politically contested category that platforms (especially Meta) have strategically inverted in EU consultations to argue access is futile. Taken together, these audits suggest that formal compliance with the DSA can coexist with substantive obstruction, a point Giglietto2026-855a54cb and de-Vreese2026-zx explicitly theorize as evidence that regulation is beginning to bite precisely because platforms are resisting it.
CrowdTangle’s shutdown and the search for alternatives
The CrowdTangle closure functions as this literature’s recurring trauma. Cullen2026-cb documents researchers’ lived experience of its constraints before shutdown; Bruns2026-pn narrates its replacement by the more restrictive Meta Content Library, deleted monthly and stripped of qualitative tooling; Lukito2026-nb catalogs it as one of several unstable Meta access regimes (Social Science One, CrowdTangle, FORT, MCL) whose cycling instability exemplifies “independence by permission.” Yet the same infrastructure, imperfect as it is, enables new empirical work: Giglietto2022-b30e8b4e uses the Social Science One URL Shares Dataset to expose an artifact of Meta’s anonymization threshold that confounds cross-country comparison, while Giglietto2026-632ef967 and Giglietto2025-1765bb4f use the Meta Content Library itself to demonstrate that Facebook actively curates visibility — dampening partisan and amplifying “quality” journalism reach in ways that track known governance interventions like “break the glass,” and showing that Meta’s political-content-reduction policy produced a documented, undisclosed 72% reach collapse for Italian MPs that asymmetrically favored extremist accounts able to compensate through volume. These papers model a cautiously optimistic register: DSA-enabled infrastructure, however flawed, permits exactly the kind of longitudinal accountability research that was impossible during the CrowdTangle era, and Giglietto2025-ed60bc90 surveys this evolving toolset directly. Zheng2026-bi offers a complementary, platform-independent solution — random-sampling tools (TubeStats, TokStats) built to bypass official APIs entirely and supply denominators that sanctioned access cannot.
Industry capture of the evidence base
A second major strand interrogates not access mechanics but epistemic independence. Bak-Coleman2026-mk and Bak-Coleman2025-pm provide the most systematic indictments: roughly half of high-profile social media research in top journals carries undisclosed industry ties, concentrated among a small recurring cohort of scientists who also dominate editorial gatekeeping, producing an estimated ~80% “industrial saturation” of the field. Heiss2026-qv situates this within cross-industry comparison (tobacco, pharma, food), arguing platform research is uniquely vulnerable because platforms alone control the data needed to study them. Munger2025-cz delivers the sharpest single-case critique, arguing the Meta2020 partnership — despite unprecedented methodological ambition — was structurally incapable of producing externally valid, temporally durable knowledge, and that its null findings conveniently served Meta’s interests via underpowered, over-conservative pre-registration. Allen2025-ot and Bechmann2026-dr respond methodologically and conceptually, respectively: the former championing platform-independent experimental tools (browser-extension LLM reranking) as an alternative to platform-cooperative designs, the latter urging a wholesale shift away from causal-effects paradigms toward theorizing platform collective behavior as democratic infrastructure. Brady2026-ln, run independently on Bluesky, exemplifies this alternative: a registered-report field experiment testing custom feed algorithms without platform gatekeeping, finding engagement-based feeds distort perceived social norms around partisan animosity even where they don’t move attitudes directly — a finding that both extends and complicates the Meta2020 null results that Gauthier2026-iq separately reconciles by showing X’s algorithm shifts attitudes conservative and asymmetrically, via a following-persistence mechanism the earlier Facebook studies missed.
Platform self-narration and the discourse of governance
Several papers examine how platforms rhetorically manage their own accountability. Gillespie2010-as supplies the foundational move — showing “platform” itself is a strategic discursive category enabling companies to claim neutrality while making consequential curatorial choices — a lineage De2026-ld updates via the concept of “changecraft,” documenting how Meta, TikTok, YouTube, and X narrate policy pivots as continuity rather than rupture. Hurcombe2025-cs dissects Meta’s Newsroom specifically, showing its “inauthenticity,” “political advertising,” “technological solutions,” and “enforcement” frames externalize responsibility for platform harms while covertly shaping regulatory language (e.g., “coordinated inauthentic behaviour”) that later diffuses into EU and Australian policy. Gillespie2022-jx extends this critique to reduction/demotion as an underexamined, largely invisible complement to removal-based moderation, one that concentrates curatorial power while evading the transparency regimes built for takedowns. Cazzamatta2026-lo and Farkas2026-lr turn to the fact-checking ecosystem’s response to this discourse: the former shows Zuckerberg’s 2025 framing of fact-checkers as censors is empirically unfounded (removal occurs in only ~30% of cases), while the latter documents how European fact-checkers rhetorically defend their platform dependencies through differentiation and transcendence — an uneasy accommodation now destabilized by Meta’s abandonment of third-party fact-checking. Holt2026-zq complicates the Community-Notes-as-successor narrative empirically, showing user “False News” reports mostly flag polarized political disagreement rather than falsity, while Renault2025-uh and Bouchaud2026-np show the crowdsourced bridging algorithm itself reproduces partisan asymmetries and structurally undermoderates precisely the polarizing content most consequential for elections.
Algorithmic power made visible through audit
A distinct empirical register treats recommender systems themselves as objects of forensic study, often explicitly framed as contributions to DSA-era systemic-risk assessment. Bouchaud2026-lr reconstructs X’s embedding space to show it inadvertently encodes users’ ideology with striking fidelity, raising GDPR/DSA profiling tensions and proposing a debiasing method. Efstratiou2025-gs and McNally2025-dn both push back against “black box” fatalism: the former shows apparent right-leaning amplification on early Musk-era Twitter is better explained by proximity to Musk and agitation than partisanship per se, while the latter demonstrates Facebook’s News Feed algorithm produces detectable, lagged, section-specific engagement effects — evidence, they argue, that systemic algorithmic auditing under Article 40(4) is technically feasible. Karo2026-dn and Rieder2026-pp extend audit logic to extremist and ambient-ideological content, documenting how TikTok’s and YouTube’s moderation architectures are structurally outpaced by their own recommendation systems, whether through jihadist “everyday extremism” evading keyword filters or Andrew Tate’s diffuse “Tate-space” persisting post-deplatforming. Votta2025-xz and Inacio-da-Silva2026-zf apply similar audit logic to political advertising delivery and irregular electoral ads, respectively, underscoring that opacity in ad systems remains a persistent governance gap even as content-level transparency improves.
Regulatory theory, comparative governance, and the democratic stakes
Zooming out, Crosset2026-mq offers the collection’s most systematic comparative-legal analysis, identifying three converging “circulation regimes” (free circulation, patrolling, flow optimisation) across US and EU legislation, all bureaucratizing security through transparency and auditing rather than direct content control — a frame that helps explain why DSA implementation looks the way Peters2026-mo, Ahuja2025-ku, and Tonneau2025-bv describe it. van-Dijck2018-up supplies deep theoretical grounding for the entire collection’s normative stakes, arguing platforms are not neutral infrastructure but active constructors of contested public value — a claim Ventura2026-yc and Bruns2026-yv validate empirically and historically, respectively: Brazil’s X ban produced durable rightward partisan sorting rather than restored pluralism, while Twitter’s broader decline shows that neither state bans nor successor platforms (Mastodon, Threads, Bluesky) easily reconstruct healthy public debate once governance and moderation fail. Donovan2025-ws and Lewandowsky2026-ob situate this decline within a longer history of platforms retreating from trust-and-safety investment after 2021, a retreat Moran2025-qn documents from inside the profession itself, and connect it explicitly to democratic backsliding. Boyd2026-op and Swartz2026-zb press this further into ontological territory, arguing that “social media” itself may be the wrong category — having become “parasocial media” organized around scam logics and one-directional influencer attention rather than reciprocal sociality — a reframing with direct consequences for what governance and data-access regimes should even be designed to observe.
Election research as the pressure point
Elections research crystallizes nearly every tension above into a single applied domain. Philipp2026-tl, Lukito2026-nb, and Schulte2026-df each use concrete election cases (EU elections generally, cross-national comparison, the 2025 German federal election) to show that multi-platform, cross-temporal electoral analysis remains infrastructurally underdeveloped despite DSA-era legal gains — fragmented APIs, inconsistent eligibility, and platform-specific biases (e.g., TikTok’s “politainment” tone-content mismatch) mean researchers still cannot reliably compare platforms or replicate findings across election cycles. Lukito2026-il shows how right-leaning outlets exploit this same fragmentation strategically, curating distinct content for Twitter versus Truth Social. Giglietto2020-6278a4aa anticipates much of this agenda pre-DSA, identifying micro-targeting, blurred organizational boundaries, and strategic amplification as three enduring measurement challenges that regulated data access has only partially resolved.
Legal risk, expert synthesis, and the road ahead
Underlying all empirical and regulatory work is the question of who can safely do it. Park2026-tr documents the chilling effects of overbroad computer-crime law on researchers who must scrape or otherwise circumvent restricted access, a risk Freelon2018-ao flagged as inherent to the post-API condition, and Pierri2026-ib treats as an active policy concern for Article 40 implementation, including a coming blind spot around LLM-based systemic risk. Two Delphi-style and multi-author syntheses attempt to draw the threads together prescriptively: Mahl2026-hc surveys 47 experts to rank platform governance and journalism-ecosystem interventions above individual-level fixes for misinformation, while Goldberg2026-eb looks forward toward AI-augmented bridging, moderation, and deliberation systems as a possible “second paradigm shift” for digital public squares — one that inherits, rather than resolves, the data-access and accountability dilemmas this entire body of work has mapped. Richter2026-bt, Schiffrin_undated-gi, Vincent_undated-re, and Graham2025-gp round out the picture by tracking how stakeholder composition, deepfake fraud, ongoing disinformation measurement, and propaganda’s exploitation of “truth infrastructure” continue to evolve as the regulatory and research landscape struggles to keep pace.