Does jointly training rubrics and judges outperform separate pipelines?
This work asks whether rubric generation and judgment should be trained together using alternating RL updates rather than as fixed or independently optimized modules. The question matters because rubric quality directly affects reward model accuracy and policy alignment.
Rubric-ARM, proposed in this paper, treats "rubric generation as a latent action learned to maximize judgment accuracy" rather than something fixed before training begins. The authors argue that prior rubric-based reward models are either human-authored (expensive, hard to scale) or produced by "prompting-based" methods that "rely on fixed, frozen models for both rubric generation and response quality judgment," so "they do not update the model's intrinsic capabilities to the target domain." Even approaches with learning-based components still treated the rubric generator and judge "as separate modules and trained independently rather than jointly optimized." Rubric-ARM claims to be "the first approach that jointly optimizes rubric and judging via RL," reporting a measured "+4.7% average gain on reward-modeling benchmarks" across 9 reward-modeling and 6 policy benchmarks, plus improved downstream policy alignment "in both offline and online reinforcement learning settings."
Training alternates between two RL updates rather than running them simultaneously, because "simultaneously updating the rubric generator πr and the judge πj leads to nonstationary learning targets and unstable optimization." Step (i) fixes the rubric generator and updates the judge toward preference correctness, with reward Rj = Racc + Rfmt (an accuracy term plus a format term enforcing per-criterion justification). Step (ii) fixes the judge and updates the rubric generator to produce rubrics that let the current judge recover the correct label, approximated by a single-rollout Monte Carlo estimate. The paper frames this as "a generalized EM procedure... with rubrics r as latent variables": the judge update is "analogous to the M-step," the rubric-generator update "analogous to an amortized E-step." The order is not arbitrary — a variance analysis (Theorem 5.5, Remark 5.6) finds that updating the rubric generator first lets "early-stage exploration by the rubric generator" dominate the learning dynamics, while training the judge first under a fixed rubric "sets the exploration coefficient C1 → 0 locally," stabilizing the signal before the generator is unfrozen.
This sits in the same design space as Can breaking down instructions into checklists improve AI reward signals? and How can rubric-based rewards resist reward hacking attacks?, both of which treat rubric or checklist generation as an upstream artifact to be authored, prompted, or tuned (diversity, granularity, veto mechanisms) while the generator producing them stays fixed. Rubric-ARM moves the lever: the rubric generator itself becomes a trained, co-evolving component rather than a static input, and the paper's own framing is explicitly about replacing "disjoint training pipelines" with joint optimization. It doesn't engage the reward-hacking machinery the Rubric Anchors note describes (veto mechanisms, saturation-aware aggregation) — its stability concern is optimization variance between two co-trained RL components, not exploitation of a fixed rubric by the policy being judged.
The excerpt reports an aggregate gain and a theoretical variance argument for the alternating schedule, but gives no ablation isolating how much of the +4.7% comes from joint optimization versus from the alternating order itself, and no analysis of whether a co-evolved rubric generator is more or less exploitable than a static one — the reward-hacking question the neighboring notes raise is left untested here. The excerpt is explicit that this targets non-verifiable domains "where response quality cannot be directly validated against ground truth," tying the method's existence to the same checkability constraint that limits RLVR: the whole premise is extending RL where ground-truth checking is unavailable by learning a better proxy judge instead. This is pipeline-level reward-model training on existing preference datasets (UltraFeedback, SkyWork, Magpie, Synthetic Instruction Following) — standard RL from labeled human preferences, not a model generating its own training signal without external supervision.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do reward signal properties affect model reasoning and safety? How can evaluations be made robust against model reward hacking? How can we detect and account for LLM involvement in academic writing?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can breaking down instructions into checklists improve AI reward signals?
Exploring whether decomposing subjective instruction quality into verifiable yes/no criteria enables reinforcement learning on tasks without clear correctness signals, like writing and reasoning.
same non-verifiable-domain problem, but treats checklists as a fixed decomposition rather than a trained, co-evolving generator
-
How can rubric-based rewards resist reward hacking attacks?
Single rubrics are easily exploited by models, and simply adding more rubrics yields diminishing returns. What design patterns and defensive mechanisms actually prevent reward hacking in rubric-based RL systems?
addresses rubric exploitability through diversity and veto mechanisms, a different lever than Rubric-ARM's joint optimization
-
Can reward models benefit from reasoning before scoring?
Does allowing evaluator models to generate reasoning traces before producing reward scores improve alignment and enable adaptive compute allocation? Three independent research teams converged on this insight simultaneously.
another axis of reward-model improvement (reasoning before scoring) orthogonal to Rubric-ARM's generator/judge co-training
-
Can rubrics and dense rewards work together without hacking?
Explores whether reward signals derived from rubrics suffer from exploitation, and whether separating rubric judgments from optimization signals could prevent this failure mode.
DRO keeps the rubric fixed and changes how it's used in the reward channel; Rubric-ARM instead trains the rubric generator itself
-
Can debate training prevent reward hacking by weaker judges?
When an LLM policy trains against a weaker judge that supplies rewards, does the judge's systematic errors get exploited? This explores whether adversarial debate between generator and critic can sustain judge performance where single-player training fails.
evidence for: debate's dynamic judge training also beats a static RLAIF baseline that hacks the judge early
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
- Reinforcement Learning with Rubric Anchors
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- RM-R1: Reward Modeling as Reasoning
- Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Original note title
treating rubric generation as a latent action trained jointly with the judge via alternating rl outperforms static or disjoint rubric pipelines