Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
Standard reward models typically predict scalar scores that fail to capture the multifaceted nature of response quality in non-verifiable domains, such as creative writing or open-ended instruction following. To address this limitation, we propose Rubric-ARM, a framework that jointly optimizes a rubric generator and a judge using reinforcement learning from preference feedback. Unlike existing methods that rely on static rubrics or disjoint training pipelines, our approach treats rubric generation as a latent action learned to maximize judgment accuracy. We introduce an alternating optimization strategy to mitigate the non-stationarity of simultaneous updates, providing theoretical analysis that demonstrates how this schedule reduces gradient variance during training. Extensive experiments show that Rubric- ARM achieves strong performance among baselines on multiple benchmarks and significantly improves downstream policy alignment in both offline and online reinforcement learning settings.
Introduction. Reward modeling serves as the compass for aligning large language models (LLMs) with human intents, typically by generating a scalar score or preference label to predict human preferences (Stiennon et al., 2020, Wang et al., 2024a). However, in complex non-verifiable domain, such as creative writing or open-ended instruction following, these scalar or pairwise judgments often fail to capture the multifaceted nature of response quality (Ying et al., 2025). To address this limitation, recent advancements have shifted toward rubric-based reward modeling, where models explicitly generate structured criteria to ground their judgments (Gunjal et al., 2026, Liu et al., 2025a, Pathak et al., 2025). By decomposing evaluation into interpretable dimensions, rubric-based models offer transparency and improve generalization across prompt-specific evaluation axes.
Central to rubric-based evaluation is the availability of high-quality rubrics. To ensure rubric quality, earlier work has primarily relied on human-authored rubrics, which are expensive to produce and difficult to scale to large datasets (Arora et al., 2025). More recent approaches seek to automate rubric construction using LLMs (Viswanathan et al., 2025, Gunjal et al., 2026); however, these methods are largely prompting-based and rely on fixed, frozen models for both rubric generation and response quality judgment. Consequently, they do not update the model’s intrinsic capabilities to the target domain or the underlying preference distribution, limiting their ability to generate in-domain, preference-aligned rubrics. Moreover, even when learning-based components are introduced (Liu et al., 2025a, Rezaei et al., 2025), the rubric generator and the judge are treated as separate modules and trained independently rather than jointly optimized. This decoupled training pipeline prevents deeper integration between rubric construction and judgment, leading to suboptimal evaluation signals. Designing effective rubricbased reward models are still challenging.
In this work, we propose Rubric-ARM, an end-to-end framework that jointly optimizes the rubric generator and the judge via alternating reinforcement learning (RL), enabling the two components to co-evolve and mutually reinforce one another during training. We formulate rubrics as latent actions that guide the reward model in recovering the underlying preference signal, and posit that improved rubric generation directly leads to more accurate preference predictions. To ensure stable joint optimization, Rubric-ARM employs an alternating training strategy that decouples the learning dynamics while preserving a shared objective. Training alternates between (i) optimizing the reward model with a fixed rubric generator to align with target preference labels, and (ii) optimizing the rubric generator with a fixed reward model to produce discriminative rubrics that maximize prediction accuracy.
A key challenge of the alternating RL is the instability caused by simultaneous updates to both components. Our analysis reveals that early-stage exploration by the rubric generator can dominate the learning dynamics. To mitigate this, we first stabilize the reward model under fixed rubrics before optimizing the rubric generator. This alternating schedule reduces variance and ensures robust optimization.
Our contributions can be summarized as follows:
• We develop Rubric-ARM, a rubric-based reward model to produce high-quality rubrics and precise judgments. To the best of our knowledge, this is the first approach that jointly optimizes rubric and judging via RL. • We introduce an alternating RL training algorithm that couples the rubric generator and judge through a shared correctness objective, enabling mutual improvement while stabilizing optimization. • We evaluate Rubric-ARM across diverse alignment settings (9 reward modeling and 6 policy benchmarks). Rubric-ARM outperforms strong reasoning-based judges and prior rubric-based reward models, achieving a +4.7% average gain on reward-modeling benchmarks, and consistently improves downstream policy post-training when used as the reward signal.
Related work. LLM-based Reward and Judge Models. While Zheng et al. (2023) established the foundational utility of LLM-based judges. Subsequent research expanded the scope of reasoning to include chain-of-thoughts (Zhang et al., 2025), self-critiques (Ankner et al., 2024, Yu et al., 2025b, Mahan et al., 2024) or plan evaluations strategically (Saha et al., 2025). Liu et al. (2025c) explore inference-time reasoning for generative reward models. Recent studies (Chen et al., 2025, 2026, Whitehouse et al., 2026, Guo et al., 2025, Hong et al., 2025, Xu et al., 2026) leverage online RL to directly incentivize detailed reasoning, aiming to mitigate bias and enhance the accuracy of pointwise and pairwise scoring.
Rubrics-based Reward Models. Recently, rubric-based approaches have emerged as a promising direction for LLM evaluation (Arora et al., 2025, Hashemi et al., 2024, Pathak et al., 2025, Akyürek et al., 2025), alignment (Viswanathan et al., 2025, Zhang et al., 2026), and reasoning (Gunjal et al., 2026, Zhou et al., 2025, Huang et al., 2025). However, a unique challenge lies in generating high-quality rubrics at scale. To address this, Li et al. (2026), Liu et al. (2025a), Xie et al. (2025) extract rubrics from pairwise comparison signals, while Rezaei et al. (2025), Zhang et al. (2026), Shao et al. (2025) dynamically generate rubrics by leveraging policy model outputs in an online setting.
Method. We study rubric-based reward modeling in non-verifiable domains, where response quality cannot be directly validated against ground truth. The rubric-based reward model contains two parts, namely rubric generator and judge. The key components of Rubric-ARM are described as follows.
Rubrics. We define a rubric as a structured set of evaluation criteria conditioned on a prompt. Formally, let x denote a prompt, a rubric r(x) = {ci}k i=1 consists of k criteria, where each ci specifies a distinct aspect of response quality (e.g., factual correctness, tone, or presentation).
In non-verifiable domains, supervision is limited to pairwise preference feedback and rubrics are not directly observed. Simultaneously updating the rubric generator πr and the judge πj leads to nonstationary learning targets and unstable optimization. As shown in Figure 1, Rubric-ARM addresses this challenge using an alternating RL scheme that decouples the updates of two components.
We equip both πj and πr with basic rubric generation and judging capabilities via leveraging open-source datasets. Following the prior work (Liu et al., 2025a), we fine-tune on synthetic rubrics and judge trajectories derived from open-source datasets including UltraFeedback (Cui et al., 2024), SkyWork (Liu et al., 2024), Magpie (Xu et al., 2025b), and Synthetic Instruction Following (Lambert et al., 2025a). Both πr(r ∣x; θr) and πj(c, o ∣x, y(1), y(2), r; θj) are trained with the standard next-token prediction objective.
Stage I (SFT) warm-starts the rubric generator πr and judge πj by imitating synthetic rubric generation and judging trajectories, but optimizes the two components independently and does not directly target preference correctness. We therefore optimize both components using alternating reinforcement learning. Specifically, training switches between (i) improving the judge with a fixed rubric generator and (ii) improving the rubric generator with a fixed judge, providing each component with a clearer learning signal while preserving the same end objective R(o, o⋆).
(i) RL for Judge πj with the current πr. With the rubric generator parameters θr held fixed, we update θj to improve preference correctness under rubrics sampled from πr:
In practice, we use a shaped reward that augments the final correctness signal Racc = I[o = o⋆ i ] with format-based reward Rfmt that enforces valid judging trajectories (i.e., addressing each rubric criterion with per-criterion explanations, followed by an overall justification and a final decision). The final reward for the judge πj is Rj = Racc + Rfmt.
(ii) RL for Rubric Generator πr with the current πj. With the judge parameters θj fixed, we update θr to prefer rubrics that lead the current judge to make correct decisions. Concretely, we maximize the preference correctness under rubrics drawn from πr as:
Intuitively, πr learns to generate criteria that are discriminative for the given prompt and usable by the judge to recover the dataset preference. In practice, we approximate the expectation with a single rollout by greedy decoding (t = 0), i.e., we generate one judging trajectory (c, o) per rubric and use the Monte Carlo estimate Rr = I[o = o⋆]. (8) Optimization (alternating RL). Rubric-ARM alternates between optimizing Eq. 5 and 7. At iteration t, we run:
Connection to EM Algorithm. Our alternating optimization can be viewed as a generalized EM procedure (Dempster et al., 1977) with rubrics r as latent variables. For each preference instance (x, y(1), y(2), o⋆), the judge defines a conditional model pθj(o⋆∣x, y(1), y(2), r), while the rubric generator πr(r ∣x; θr) acts as an amortized variational distribution over the latent rubric (Agrawal and Domke, 2021). With πr fixed, updating πj maximizes the expected correctness (or log-likelihood) under sampled rubrics, analogous to the M-step. With πj fixed, updating πr increases probability mass on rubrics that make the current judge more likely to recover o⋆, analogous to an amortized E-step.
Discussion. Remark 5.6 (Implication for Training Stability). The variance gap derived in Theorem 5.5 justifies the proposed training schedule (We first train the judge, then train the rubric generator, and subsequently perform alternating training following this sequence.) by highlighting a critical trade-off in Signal-to-Noise Ratio (SNR). The strictly higher variance in Strategy B implies that generator updates are dominated by exploration stochasticity rather than the true gradient direction, risking optimization instability. In contrast, Strategy A acts as a variance reduction mechanism: by fixing the rubric, it effectively sets the exploration coefficient C1 →0 locally, isolating the judge from structural noise and providing a stable target for effective learning.
Conclusion. In this work, we propose Rubric-ARM, a novel framework for reward modeling in non-verifiable LLM post-training. Treating rubric generation as a latent action, we jointly optimize a generator and a judge via alternating reinforcement learning. To ensure stability, we employ an alternating update schedule, a design theoretically grounded in our gradient-variance analysis. Empirically, Rubric-ARM achieves 4.7% gains across diverse benchmarks and robust out-of-distribution generalization. It also delivers superior supervision for policy alignment in both offline and online RL settings, showing Rubric-ARM offers a more reliable reward signal than static approaches.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do reward signal properties affect model reasoning and safety?- How do checklist-based rewards decompose complex judgment into verifiable criteria?
- Do spurious rewards activate reasoning without teaching new skills?
- Can log-likelihood loss combined with binary rewards achieve calibration?
- Why do reward models learn surface-level shortcuts instead of genuine quality assessment?
- Do outcome-only reward signals miss step-level errors that compound later?
- Can critic model trios evaluate reasoning quality more reliably than outcome rewards alone?
- Can intrinsic reward signals extend beyond mathematics to medicine and law?
- How do semantic reward shaping approaches compare to full critique models?
- What information do numerical rewards fail to provide for reasoning tasks?
- When should action deliberation trigger during reasoning steps?
- Do explicit reasoning chains improve or harm performance on complex judgment tasks?
- How do outcome and process rewards differ in their treatment of intermediate steps?
- Can solution traces substitute for process-level reward signals in math reasoning?
- What makes process-level supervision better than outcome-only reward signals?
- Why do process reward models need human annotation while MCTS intermediate nodes don't?