STaR trains AI by keeping only reasoning that got the right answer — but what grades it when there's no answer key?
How does STaR's correctness filtering extend to tasks without ground truth answers?
This explores what happens to STaR's self-teaching recipe (keep only the reasoning that led to a correct answer, then train on it) when there is no answer key to check against, as in open-ended writing, advice, or general-knowledge reasoning.
This explores what replaces the answer key once STaR's simple self-teaching trick moves into domains where 'correct' can't be checked automatically. The original idea is very simple. A model writes out its reasoning, you keep only the attempts that reached the right answer, and you fine-tune on those. That alone pushed a small model on CommonsenseQA close to models 30 times its size Can models improve by filtering only on answer correctness?. The whole method depends on that filter, though. Take away ground truth and you need some other way to decide which self-generated reasoning deserves to be kept. The corpus has no single paper titled 'STaR without labels.' What it does have is several separate attempts to build that replacement filter, and each one shows a different way it can fail.
The closest descendant gives up on judging correctness and switches to a softer question. VeriFree doesn't ask 'is this answer right?' It asks 'given this reasoning, how likely is the model to produce the reference answer?' and uses that probability as both the reward and the training weight Can reasoning improvement work without answer verification?. It matches verifier-based methods on broad benchmarks like MMLU-Pro and GPQA. The catch is that it still needs a reference answer. What it removes is the need for an exact-match checker. So the first lesson is that 'no ground truth' usually means 'no cheap way to check,' not 'no target at all.'
The obvious next step is to let the model grade its own work, and this is where things go wrong. Models systematically over-trust answers they generated themselves, because a high-probability output simply feels right when they evaluate it Why do models trust their own generated answers?. More generally, LLMs are good at proposing candidates but poor at estimating how good those candidates actually are Can language models reliably judge their own candidate quality?. A STaR loop filtered by self-judgment can therefore reinforce whatever the model already believed. This matters because fluent, confident wrong answers already hide inside good-looking aggregate scores Why do confident wrong answers hide in standard accuracy metrics?. A self-training loop can quietly multiply exactly those errors.
The more promising replacements come from outside the training literature. One finding treats verification as its own scaling axis: judges get much more accurate when you break the criteria into parts, score more finely, and repeat the evaluation, all without retraining Can verification accuracy scale without training models?. Agent-style judges that actively gather evidence cut evaluation drift roughly 100-fold compared with a plain LLM judge Can agents evaluate AI outputs more reliably than language models?. The most unexpected parallel is in retrieval systems. One RAG design writes its own answers back into its knowledge base, but only after they pass entailment, source-attribution, and novelty checks Can RAG systems safely learn from their own generated answers?. That is the same pattern as STaR: the system generates, filters, and keeps what passes. The filter is just grounding in evidence instead of answer-matching.
What you might not have expected to learn: once ground truth is gone, the filter becomes the real thing being trained, and its errors compound with every round. Either it shrinks to a softer probability signal that still leans on references, or it has to become an evidence-checking system that is often more elaborate than the model it supervises.
Sources 8 notes
STaR demonstrates that self-generated rationales filtered exclusively by answer correctness improve reasoning performance significantly. On CommonsenseQA, this correctness-filtered approach achieved 72.5% accuracy, outperforming direct answer fine-tuning and closing the gap with models 30 times larger.
VeriFree bypasses answer verification entirely by using the conditional probability of reference answers given generated reasoning traces as both reward signal and training weight. This approach matches or surpasses verifier-based methods on MMLU-Pro, GPQA, and SuperGPQA without rule-based or model-based verifiers.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.
Medical triage, legal interpretation, and financial planning show a consistent pattern: surface heuristics conflict with unstated constraints, producing fluent confident errors that concentrate in rare cases where harm occurs. Aggregate accuracy masks these failures because overall performance looks strong.
Show all 8 sources
Research shows verification accuracy improves independently via score granularity, repeated evaluation, and criteria decomposition—all deployable at inference without retraining. This reframes weak verifiers as under-scaled rather than fundamentally limited.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Systems can add generated answers to their retrieval corpus when outputs pass entailment verification, source attribution checks, and novelty detection. This prevents hallucinations from polluting future retrievals while allowing genuine knowledge accumulation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Large Language Models Cannot Self-Correct Reasoning Yet
- Reinforcing General Reasoning without Verifiers
- Self-reflective Uncertainties: Do LLMs Know Their Internal Answer Distribution?
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- UR2: Unify RAG and Reasoning through Reinforcement Learning