Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

Paper · arXiv 2609.15624 · Published September 14, 2026
Workplace Applications

AI-literacy measures now include self-report questionnaires, objective tests, and instruments for critical oversight and reliance. Their different targets complicate assessment of competent generative-AI use in work settings. We conducted a structured, seeded evidence review anchored in the 2024 COSMIN-based review of AI-literacy scales, with a targeted update through 17 August 2026. The synthesis covers 24 focal empirical publications plus the prior review and organizes reported measurement content into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. An exploratory meta-analysis of three directly reported subjective-objective correlations, all from one research program, produced a REML pooled r = .055 with a Hartung-Knapp 95% confidence interval of [-.047, .156]. The model uses a combined reported N = 2,765; the largest study has an unresolved discrepancy between its reported correlation and p-value, so its weighting requires caution. Adding a synthetic mean of 12 cross-factor correlations from a fourth study yielded r = .079, 95% CI [-.025, .181]. This composite sensitivity addresses a broader comparison than the direct-effect analysis.

Introduction. Competent AI use now includes decisions that earlier AI-literacy frameworks did not need to test. Those frameworks asked whether a person could recognize AI, understand broad mechanisms, use applications, evaluate outputs, and reason about ethics [Long and Magerko, 2020, Ng et al., 2021]. Those domains remain useful, but tool-using assistants add decisions about access, state changes, execution, and evidence. A person can answer conceptual questions about large language models and still make poor operational decisions. Examples include accepting an unsupported claim because the prose sounds convincing, granting a tool more access than a task requires, allowing execution to continue after the plan has changed, repairing a corrupted session instead of returning to a known-good state, or treating an agent’s completion report as proof that required checks ran. These behaviors sit between AI literacy, critical thinking, reliance, human factors, and professional task competence. Measurement research has addressed each neighboring area in part, but the boundaries between them remain loose.

Discussion / Conclusion. The reviewed instruments assess different components of competent GenAI use. Objective tests such as AICOS-S and GLAT assess demonstrated foundation knowledge; self-reports assess perceptions, confidence, and reported practice. The exploratory primary analysis of three direct subjective-objective correlations yielded r = .055, with a Hartung-Knapp interval that includes zero. All three effects come from one research program, and the largest study’s weighting depends on an unresolved reporting discrepancy. Adding a cross-factor composite from another study gives r = .079, also with an interval that includes zero. These results provide no basis for substituting self-ratings for performance scores, but do not establish a population correlation or a workplace pass threshold. Other instruments address verification, reflective oversight, reliance, trust, and dependency.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI fluency substitute for verifiable accuracy in human judgment? Can AI-generated outputs constitute genuine knowledge or valid claims? How does AI-generated content transformation affect public discourse quality? How can humans calibrate appropriate trust in AI systems? How does AI adoption affect human skill development and labor equality? How do we evaluate AI systems when user perception misleads actual performance? Why does verification consistently lag behind AI generation? How does AI assistance affect human cognitive development and reasoning autonomy? Does AI text rewriting systematically distort writer intent and preference?