Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use
AI-literacy measures now include self-report questionnaires, objective tests, and instruments for critical oversight and reliance. Their different targets complicate assessment of competent generative-AI use in work settings. We conducted a structured, seeded evidence review anchored in the 2024 COSMIN-based review of AI-literacy scales, with a targeted update through 17 August 2026. The synthesis covers 24 focal empirical publications plus the prior review and organizes reported measurement content into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. An exploratory meta-analysis of three directly reported subjective-objective correlations, all from one research program, produced a REML pooled r = .055 with a Hartung-Knapp 95% confidence interval of [-.047, .156]. The model uses a combined reported N = 2,765; the largest study has an unresolved discrepancy between its reported correlation and p-value, so its weighting requires caution. Adding a synthetic mean of 12 cross-factor correlations from a fourth study yielded r = .079, 95% CI [-.025, .181]. This composite sensitivity addresses a broader comparison than the direct-effect analysis.
Introduction. Competent AI use now includes decisions that earlier AI-literacy frameworks did not need to test. Those frameworks asked whether a person could recognize AI, understand broad mechanisms, use applications, evaluate outputs, and reason about ethics [Long and Magerko, 2020, Ng et al., 2021]. Those domains remain useful, but tool-using assistants add decisions about access, state changes, execution, and evidence. A person can answer conceptual questions about large language models and still make poor operational decisions. Examples include accepting an unsupported claim because the prose sounds convincing, granting a tool more access than a task requires, allowing execution to continue after the plan has changed, repairing a corrupted session instead of returning to a known-good state, or treating an agent’s completion report as proof that required checks ran. These behaviors sit between AI literacy, critical thinking, reliance, human factors, and professional task competence. Measurement research has addressed each neighboring area in part, but the boundaries between them remain loose.
Discussion / Conclusion. The reviewed instruments assess different components of competent GenAI use. Objective tests such as AICOS-S and GLAT assess demonstrated foundation knowledge; self-reports assess perceptions, confidence, and reported practice. The exploratory primary analysis of three direct subjective-objective correlations yielded r = .055, with a Hartung-Knapp interval that includes zero. All three effects come from one research program, and the largest study’s weighting depends on an unresolved reporting discrepancy. Adding a cross-factor composite from another study gives r = .079, also with an interval that includes zero. These results provide no basis for substituting self-ratings for performance scores, but do not establish a population correlation or a workplace pass threshold. Other instruments address verification, reflective oversight, reliance, trust, and dependency.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI fluency substitute for verifiable accuracy in human judgment?- Why are less experienced thinkers more vulnerable to false AI credibility?
- Why do users interpret AI outputs through frameworks meant for human experts?
- How does AI reduce the skill gap between amateur and expert-level misuse actors?
- How does validation skill replace production skill in AI systems?
- Why do users feel more competent when their actual capability is declining?
- What mechanisms make users misattribute AI outputs as their own competence?
- Why do users believe they produced independent competence when they actually used AI assistance?
- What happens when confident language masks uncertainty in AI outputs?
- Why do people misattribute AI outputs as evidence of their own skill?
- Why does AI fluency create false impressions of expert judgment?
- Why does polished AI output feel like evidence of user skill?
- Why don't users push back when AI makes obvious mistakes about false claims?
- Can disclaimers alone prevent users from trusting AI outputs too heavily?
- Why do users trust overconfident AI outputs across different languages?
- What happens when AI-dependent workers must operate without their tools?
- How should AI systems model human resource constraints and expertise levels?
- Why do AI model updates cause genuine grief in users?
- Why do workers who debug most with AI show the lowest learning outcomes?