If your school won't say yes or no to AI on assignments, how do you decide what's fair anyway?
How should teachers make GenAI assessment decisions without institutional permission?
This explores how individual teachers can make sensible calls about generative AI in their assessments when their institution hasn't given clear rules or permission, and what the collection says about acting under that uncertainty.
This explores how a teacher can decide what to do about generative AI in assessment when no institutional rule tells them what's allowed. One caveat first: the collection has no study of teachers acting without permission. What it offers instead is a set of findings that change how the decision looks. The most useful reframe comes from interviews with 20 university teachers, which found that the GenAI assessment challenge meets all ten criteria of a 'wicked problem' Does GenAI assessment challenge fit wicked problem theory?. It has no agreed definition, no point at which it's solved, and only better or worse responses rather than right answers. That matters for a teacher waiting on permission. If the problem is wicked, a definitive institutional policy may never arrive, and the realistic job is to make judgments you can defend and keep revising them.
There's also evidence that rules may matter less than people assume. In a randomized experiment on peer review at ICML 2026, banning LLM use and allowing limited use produced almost the same scores, decisions and reviewer confidence. Meanwhile, a large share of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. Peer review isn't a classroom, but the lesson carries over: a ban or permission on paper doesn't settle what actually happens. That moves attention away from 'am I allowed?' and toward assessment design. Work on the GPQA benchmark shows that carefully written expert questions can stay hard even for skilled people with full web access Can expert-written questions resist web-assisted non-expert answering?. That suggests it's possible to design tasks that are hard to answer just by looking things up, rather than relying on detection or prohibition.
A less obvious finding concerns openness. Interviews with knowledge workers show that people often hide signs of GenAI use, partly to avoid stigma but also to signal their own expertise. That hiding cuts off the informal peer learning organizations need to adapt Why do knowledge workers hide signs of using GenAI?. For a teacher acting without clear permission, the temptation is to experiment quietly. The research suggests that sharing your reasoning with colleagues may be what helps a department work out its norms when policy hasn't caught up.
If you're thinking of using AI on your own side of assessment, the evidence splits. For writing questions, ChatGPT-generated practice items matched published textbook questions on difficulty and on how well they told stronger and weaker students apart Can AI generate assessment questions as good as human experts?, so this is a relatively low-risk place to start. Grading is riskier. AI evaluators give higher scores to answers with fake references or polished formatting regardless of content Can LLM judges be tricked without accessing their internals?. They also tend to prefer text they recognize as their own Do LLMs favor their own text because they recognize it?, and they forgive rule-breaking when they think a human wrote the work Do authorship labels change how AI judges evaluate rule violations?. In practice: using AI to help draft questions is easy to defend, but letting it grade students on its own is hard to justify, whatever your institution has or hasn't approved.
Sources 8 notes
Analysis of 20 teacher interviews at an Australian university shows the GenAI-assessment challenge matches every characteristic of wicked problems: no agreed definition, no stopping rule, only better-or-worse solutions. This explains why policy and detection tools alone fail.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
GPQA's 448 expert-vetted questions in biology, physics, and chemistry achieved 65% accuracy with PhD-level experts but only 34% with web-equipped non-experts, suggesting the benchmark resists retrieval-based solving and may serve as a scalable-oversight testbed.
Interviews with 19 knowledge workers across sectors reveal that erasing GenAI cues serves as a positive expertise signal, not only stigma avoidance. This concealment reduces informal peer knowledge-sharing and reinforces organizational cultures lacking GenAI transparency.
A controlled study of 207 respondents found ChatGPT-generated formative assessment items were statistically equivalent to published textbook questions on difficulty, discrimination, and response time using IRT methodology. Items showed no disruption to measurement validity.
Show all 8 sources
Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- The Wicked Problem of AI and Assessment
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Evaluating the psychometric properties of ChatGPT-generated questions