Matias, J. N., & Price, M. (2025). How public involvement can improve the science of AI. Proceedings of the National Academy of Sciences, 122, e2421111122. https://doi.org/10.1073/pnas.2421111122
Summary
This perspective argues that public involvement—especially engagement with lived-experience experts—can materially improve the scientific rigor of AI evaluation, not just its social or political acceptability. Matias and Price treat AI systems as sociotechnical artifacts whose reliability, performance, and security depend on social, organizational, and political layers that purely technical evaluations routinely overlook. Drawing on participatory science literature and case studies across healthcare, policing, child welfare, truth commissions, and social media, they identify five distinct stages of quantitative AI evaluation—equipoise, measurement, explanation, inference, and interpretation—where public participation strengthens the underlying science. Their core claim is that trustworthy AI depends on trustworthy science, which in turn depends on inclusive, participatory processes.
Key Contributions
- A five-part framework (equipoise, measurement, explanation, inference, interpretation) for integrating public involvement into quantitative AI evaluation.
- A bridge between qualitative/critical sociotechnical scholarship and quantitative AI evaluation practice.
- A systematic rebuttal of common scientific objections to participatory research (generality, subjectivity, reliability, cost/pace/scale).
- A typology of public-involvement models—contributory, cocreated, participatory governance—applied to AI.
- Concrete case studies showing how lived-experience expertise improved real evaluations.
- The framing of trustworthy AI as contingent on trustworthy, participatory science.
Methods
A conceptual, review-based perspective synthesizing literature on participatory science, AI evaluation, fairness/ML, and sociotechnical systems. The argument is grounded in the authors’ 15+ years across generative AI evaluation, human rights investigations, algorithm audits, predictive policing analyses, and field experiments on human–algorithm behavior. Illustrative cases include Kaiser Permanente nursing protests and organ allocation, Chicago police complaints via the Invisible Institute, the Allegheny Family Screening Tool, HRDAG’s truth-commission work in Guatemala and Colombia, and Reddit field experiments. It also engages philosophy of science, invoking Popperian falsification and equipoise from medical ethics.
Findings
- Mission-critical AI faces interlocking reliability, contextual-performance, security, and transparency challenges that purely technical evaluation cannot resolve.
- Contributory citizen-science models raise consent and worker-treatment concerns but can produce robust data when participants find meaning in contributing.
- Cocreated evaluation surfaces flaws developers miss: HRDAG/ACLU’s reanalysis of the Allegheny Family Screening Tool revealed bipartite-ranking bias against Black families hidden by the original AUC-based evaluation.
- Community labeling of Chicago police complaints uncovered allegations of sexual violation obscured by official single-category coding.
- A participatory Reddit field experiment yielded a sociotechnical theory of how human fact-checking can affect recommender algorithms.
- Equipoise lets adversarial parties commit to evidence-based processes despite competing interests.
- Community debriefing can preserve study validity by identifying mid-study platform changes and unobserved confounders.
Connections
This paper’s participatory, lived-experience-driven stance on AI evaluation directly complements work on participatory and community-grounded red-teaming and evaluation such as Gillespie2026-aa, Ng2026-og, and Jayaram2026-wd. Its emphasis on field experiments and sociotechnical inference about human–algorithm behavior also relates to empirical evaluation work like Hackenburg2026-ud and Unknown2025-qj.