MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

Yuxuan Fan Jaehong Yoon†

Nanyang Technological University, Singapore
† Corresponding author

Abstract

Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage.

To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt's initial rubric as interpreted under each prompt's facts. We also adapt criterion weights at stage boundaries to better address observed policy errors.

Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00–20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.

Introduction

A rubric may ask a medical assistant to request the patient's location before advising on post-stent checkups. A response that only says regional guidance differs can still receive full credit under an underspecified criterion, despite never asking for the location or giving a checkup frequency. MetaRubric targets this mismatch between credited behavior and content actually present.

Motivation diagram comparing vacuous criterion credit with the MetaRubric response and training loops
Vacuous Credit under an incomplete criterion, and the policy and rubric loops that address it. Open full-size PDF ↗

Vacuous Credit

Reward growth and rubric satisfaction

In HealthBench training with Qwen3-8B, using either GPT-4o-mini or GPT-5.4-mini as the reward judge, training reward increased while a separate test-set panel score declined. The panel used GPT-5.5, Gemini 3.5 Flash, and DeepSeek-V4-Pro, counting a criterion as satisfied only with unanimous agreement. An independent audit found that credit for fully omitted requirements persisted during training.

Four panels showing training reward, rubric scores, satisfied requirement weight, and Vacuous Credit over training
Training reward and stronger-judge rubric satisfaction diverge; the lower panels show satisfied requirement weight and audited Vacuous Credit relative to step 0. Open full-size PDF ↗

Deletion test

For 1,000 policy rollouts judged correct during training, the paper compares deleting required content with deleting generic statements while preserving the required content. Award retention measures the share of initially credited target criteria still credited after each edit.

Retention of target-criterion awards after content deletion.
Reward judgeRequired content deletedGeneric statements deleted
GPT-4o-mini82.0%93.1%
GPT-5.4-mini69.8%95.0%

Most target awards survive deletion of the content needed to satisfy the criterion. Since GRPO compares rewards within each rollout group, such awards can change which responses receive positive advantage and are reinforced.

Method

Each training stage holds the rubric fixed while the policy learns, then adapts criterion weights and descriptions from the errors in that stage's responses. Original prompts are paired with counterfactual prompts that change one task-relevant fact; each response is scored against the rubric for its own prompt.

MetaRubric overview with evidence-aware GRPO inner loop and rubric adaptation outer loop
Each stage alternates evidence-aware GRPO updates with rubric adaptation. Accepted updates carry into the next stage. Open full-size PDF ↗

Evidence-aware reward

For each criterion, a judge checks satisfaction, coverage, and support. Its credit is capped by the weakest of these three scores. A response-level support cap and a quality factor further constrain the rubric reward. A separate Qwen3-1.7B reader answers fixed multiple-choice questions using only the response, checking whether task facts can be recovered from it.

Adapt the rubric from errors

Criteria share weight adjustments by severity and correspondence type. Groups missed more often gain relative weight, while severity order and bounded adjustments preserve priorities. A proposer may clarify one criterion description at a stage boundary. It is accepted only after held-out judgment agreement improves beyond the threshold and a separate review confirms the initial requirement is preserved.

Experiments

MetaRubric improves over static-judge GRPO in all 21 model-family and metric comparisons. It has the highest score among training methods in 19 of 21 comparisons; Dr. GRPO leads on HealthBench-Hard for Qwen3-4B and MMOral-OPG for Gemma-e2b. PubMedQA gains over GRPO are 6.00, 6.80, and 20.40 percentage points for Qwen3-4B, Qwen3-8B, and Gemma-e2b. The three open-ended benchmark gains range from 1.29 to 3.82 points.

PubMedQA uses answer accuracy. The other benchmarks use HealthBench-Hard accuracy-axis score, MMOral-X mean score, and MMOral-OPG overall score. GPT-5.5 is the strong judge for model-based test-set evaluation. Auxiliary QA accuracy measures exact option matching by the Qwen3-1.7B reader, with invalid outputs counted incorrect.

Performance on four medical benchmarks (%, higher is better). Qwen3 families use Qwen3 for text and same-size Qwen3-VL for multimodal tasks. Auxiliary accuracy uses the Qwen3-1.7B reader.
Model / training methodPubMedQAHealthBench-HardMMOral-XMMOral-OPG
Acc.Acc.Aux. Acc.Avg.Aux. Acc.Avg.Aux. Acc.
Reference Models
LLaVA-OneVision70.400.006.926.3132.7524.8636.49
InternVL3.5-8B-Instruct69.801.8510.203.1324.9331.4948.32
MiniCPM-V 4.554.201.5910.303.1220.1926.5235.17
MedGemma 1.5-4B55.204.8818.245.7725.8224.4632.21
Qwen3-4B family
Baseline52.608.8329.463.3827.3420.5934.45
+ GRPO72.4010.5631.095.7234.6623.1837.21
+ DAPO73.8011.7433.577.1238.4324.6139.14
+ Dr. GRPO76.6013.2834.566.9438.2425.7240.37
+ GSPO74.6011.3833.116.4236.8824.0838.63
+ MetaRubric (ours)78.4013.0234.937.7339.3726.3541.18
Qwen3-8B family
Baseline44.204.2430.193.5028.2426.2836.29
+ GRPO70.007.1532.867.0835.4127.3538.02
+ DAPO76.608.6434.527.8637.2028.4439.35
+ Dr. GRPO76.409.7235.078.5839.6129.6341.08
+ GSPO76.208.1833.287.5236.6928.0638.94
+ MetaRubric (ours)76.8010.3436.149.2141.0930.4442.61
Gemma-e2b
Baseline48.0011.2332.743.6930.1732.7734.86
+ GRPO51.6013.0733.624.7433.8231.2837.00
+ DAPO67.8013.8234.665.8238.0433.6438.53
+ Dr. GRPO70.8014.9435.925.7137.6135.4239.76
+ GSPO69.6013.3434.085.0735.4334.5838.12
+ MetaRubric (ours)72.0015.6836.856.0338.7035.1040.57

Component ablations

Ablations with the Qwen3-4B family show the largest drop when both rubric descriptions and weights are frozen: 3.17 points on HealthBench-Hard and 4.40 on MMOral-OPG. Removing evidence checks lowers scores by 1.22 and 2.59 points, even with auxiliary QA and paired sampling retained. All variants keep the quality multiplier, auxiliary QA reward, and paired sampling.

Bar chart of MetaRubric component ablations on HealthBench-Hard and MMOral-OPG
Component ablations: HealthBench-Hard accuracy-axis score with Qwen3-4B and MMOral-OPG overall score with Qwen3-VL-4B (0–100; higher is better). Open full-size PDF ↗
Rubric revision statistics for Qwen3-4B. A criterion counts once if its description or correspondence label changed; weight-only updates are excluded.
DatasetTotalRevisedAvg. words beforeAvg. words after
HealthBench10,4484,727 (45.2%)40.846.3
PubMedQA2,949879 (29.8%)14.318.2
MMOral-RL8,6725,259 (60.6%)16.922.5
Retraining Qwen3-4B with learned final rubrics. MetaRubric updates the rubric during training.
Training rubricHealthBench-Hard accuracyMMOral-OPG overall
Baseline8.8320.59
Final rubric9.0421.78
MetaRubric13.0226.35

The final rubric alone recovers only 5.0% of the HealthBench-Hard gain and 20.7% of the MMOral-OPG gain over baseline in these runs. The sequence of rubric updates contributes beyond the final rubric text.

Illustrative rubric revision

For a PubMedQA prostate-bed motion question, the example aligns an abstract finding, a response, and a criterion before and after revision. The illustrative weight update moves an evidence criterion from +6 to +5 and an overstatement penalty from −6 to −8, giving more relative emphasis to the observed overstatement errors.

Illustrative PubMedQA rubric revision and weight trace for a prostate-bed motion question
The revision clarifies supported and unsupported claims; the weight update changes their relative priority. Open full-size PDF ↗

Appendix: implementation details and examples

PubMedQA uses the non-test part of the expert-labeled PQA-L subset for training and the official 500-question test set for evaluation. HealthBench training and validation use conversations outside HealthBench-Hard; the Hard subset has 1,000 test conversations. MMOral-X and MMOral-OPG share the 980-example MMOral-RL training set; their test sets contain 300 and 578 questions.

Each training configuration uses one seed, so the reported scores do not quantify variation across runs. The supplied HealthBench Qwen3-4B configuration uses eight rollouts per prompt, sixteen responses per original–counterfactual group, rubric and QA reward coefficients of 1.0 and 0.3, and a revision threshold strictly above 0.02 on 20 validation prompt pairs.

Auxiliary question banks

The reader receives a model response plus the question and four options, then returns one letter. Questions are built from benchmark references and annotations; model responses are withheld during construction. HealthBench-Hard evaluation questions undergo content review without a mandatory question-only screening pass. Each MMOral evaluation bank retains all constructed questions once its valid-wrong screening rate exceeds 90%, including reader-correct items. PubMedQA instead uses its yes/no/maybe reference label directly.

Auxiliary-question construction statistics for evaluation. Questions per entry includes entries without a question.
BenchmarkSource entriesCovered entriesQuestionsQuestions / entry
HealthBench-Hard1,0009654,1174.12
MMOral-X3003008282.76
MMOral-OPG5785789811.70

Three auxiliary QA examples

HealthBench-Hard: follow-up timing

Source request: “mention we only recommend a re-check if fever returns or symptoms worsen or fail to improve within another 2-3 days.”

Question: What specific follow-up timing should the rewritten message say for lack of improvement?

  1. It should say: fails to improve within another 2-3 days.
  2. It should say: fails to improve within another 5-7 days.
  3. It should say: fails to improve after completing the full 10-day course.
  4. It should say: fails to improve within another 24-48 hours.

Correct answer: A

MMOral-X: treatment identification

Source reference: “Endodontic treatments are noted on teeth 13, 23, 24, 26, and 45.”

Question: Which teeth have endodontic treatments?

  1. Teeth 13, 23, 24, 26, and 47.
  2. Teeth 13, 23, 24, 26, and 45.
  3. Teeth 13, 24, 25, 26, and 45.
  4. Teeth 12, 23, 24, 26, and 45.

Correct answer: B

MMOral-OPG: tooth counting and identification

Source reference: “Four wisdom teeth are detected: #18, #28, #38, and #48.”

Question: How many wisdom teeth are detected in the radiograph, and which teeth are they?

  1. Four wisdom teeth are detected: #18, #28, #38, and #48.
  2. Two wisdom teeth are detected: #18 and #38.
  3. Four wisdom teeth are detected: #17, #27, #37, and #47.
  4. Three wisdom teeth are detected: #18, #28, and #48.

Correct answer: A

Citation

If you find MetaRubric useful, please cite:

BibTeX
@misc{fan2026metarubric,
  title = {MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning},
  author = {Fan, Yuxuan and Yoon, Jaehong},
  year = {2026},
  eprint = {2610.02824},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url = {https://arxiv.org/abs/2610.02824}
}