The LLM StackFrom Silicon to Agents
Part VI — RL Infrastructure (Deep Dive)
24 min read·Updated ·▶ Run the code (Colab)

6.8 Reward Engineering, Verifiers & Sandboxes

The training signal in reinforcement learning from human feedback (RLHF) and in RL with verifiable rewards (RLVR) is just a scalar \(r \in \mathbb{R}\) returned for each model completion. Everything the model learns — its reasoning habits, its code style, its factual accuracy — is shaped by what that scalar measures. Getting the reward function right is therefore the most consequential engineering decision in your entire post-training pipeline.

This chapter is a practical deep-dive into reward engineering: how to build rule-based verifiers that check math answers or run code tests, how to sandbox arbitrary execution safely at scale, how to use a second LLM as a judge, how to combine multiple reward signals without one dominating, and how to instrument and serve the whole system at training throughput. We assume familiarity with the basic RLHF/RLVR pipeline described in RL with Verifiable Rewards (RLVR) & The Reasoning Recipe and the training loop mechanics from The Generation–Training Loop & Rollout Engines.

The Anatomy of a Reward Function

Every reward function takes a (prompt, completion) pair and optionally a reference answer or test suite, and returns a scalar. In practice they come in three families:

Family Examples Pros Cons
Rule-based verifier Math equivalence, regex match, code unit test Deterministic, zero cost at training time Only defined for tasks with ground truth
LLM-judge A frontier model (GPT-5, Claude) scoring correctness, helpfulness Works for open-ended tasks Expensive, noisy, gameable
Learned reward model Bradley-Terry model trained on preference data Smooth signal, general-purpose Reward hacking, needs fresh data

In RLVR (the paradigm that produced DeepSeek-R1 and related reasoning models), the goal is to rely on the first family as much as possible. The key insight is that for tasks with a checkable ground truth, a rule-based verifier is both cheaper and harder to hack than a learned reward model. We still need LLM judges and learned reward models for tasks like “write a good essay,” but we should use them sparingly.

Reward Signal Flow

Training Process Prompt Policy (LLM) Completion PREFERRED Rule-Based Verifier cost: cheap signal: deterministic hard to hack LLM Judge cost: expensive signal: noisy gameable Reward Model cost: GPU signal: smooth hackable Weighted Combination r_total r_total = sum(w_k * r_k) Advantage Estimation Policy Loss gradient update
Reward signal flow in the RLVR training loop. Each training step generates a completion from the policy, fans it out to three parallel verifiers running at different latencies and costs, merges their outputs into a weighted scalar, computes advantages, and updates the policy via a gradient — closing the loop.

The weighted combination step is itself a design choice (see section on multi-objective rewards below). For now, note that different reward sources run at different latencies and different costs, which matters a lot for training throughput.

Rule-Based Verifiers

Math Equivalence Checking

The canonical RLVR task is mathematical reasoning: generate a chain-of-thought and a boxed final answer, then check if the answer is mathematically equivalent to the ground-truth. “Mathematically equivalent” is harder than string equality.

Consider: 1/2, 0.5, \frac{1}{2}, 50\%, and 0.50 all represent the same value. A naive string comparison gives sparse reward and introduces arbitrary formatting bias. We need symbolic or numeric normalization.

The same value, many costumes 1/2 underlying value "1/2" "0.5" "\frac{1}{2}" "50%" "0.50" string == FALSE equal value, unequal strings -> naive == gives sparse, biased reward math_equivalent(pred, gold): the decision ladder 1. Normalize strip \boxed{...}, drop $, collapse whitespace -> exact string match? yes 2. Numeric both parse as numbers (fractions, percent) -> compare with relative tolerance ~10^-6 yes 3. Symbolic parse LaTeX with sympy -> simplify(pred - gold) == 0 yes 4. Numeric fallback evaluate expr to float, N(expr) -> tolerance compare yes no no no EQUIVALENT (reward = 1) all four fail NOT equivalent (reward = 0) A buggy normalizer that rejects \frac{1}{2} == 0.5 is a false negative -> a phantom ceiling on reward (a reward-hacking surface).
Mathematically equivalent values wear many surface forms, so a rule-based math verifier cannot check reward with string equality. math_equivalent() instead runs a four-rung cascade — normalize, numeric compare, symbolic simplification, then numeric fallback — short-circuiting to EQUIVALENT the moment any rung matches, and only falling through to reward 0 if all four fail; a bug in any early rung silently caps the reward the policy can ever receive.
r"""
math_verifier.py — robust math answer checker for RL training.

Handles:
  - Fraction strings: "1/2", "3 1/4" (mixed numbers)
  - LaTeX: r"\frac{3}{4}", r"\sqrt{2}", r"2^{10}"
  - Percentages: "50%" -> 0.5 comparison
  - Sets/tuples: "{1, 2}" == "{2, 1}"
  - Symbolic via sympy with numeric fallback
"""

import re
import math
from fractions import Fraction
from typing import Optional, Union
import sympy
from sympy.parsing.latex import parse_latex
from sympy import simplify, N


def match_braces(s: str, open_idx: int) -> Optional[int]:
    """
    Given the index of a '{' in s, return the index of its matching '}'.

    Brace matching (not regex) is mandatory here: competition answers nest
    braces, so a lazy regex over a boxed group stops at the FIRST '}' and
    turns \\boxed{\\frac{1}{2}} into the garbage string '\\frac{1'. That
    single bug is a common cause of a phantom accuracy ceiling in RLVR runs.
    """
    depth = 0
    for i in range(open_idx, len(s)):
        if s[i] == '{':
            depth += 1
        elif s[i] == '}':
            depth -= 1
            if depth == 0:
                return i
    return None  # unbalanced (truncated generation)


def strip_boxed(s: str) -> str:
    """Replace the first \\boxed{...} with its (brace-balanced) contents."""
    m = re.search(r'\\boxed\s*\{', s)
    if not m:
        return s
    open_idx = m.end() - 1
    close_idx = match_braces(s, open_idx)
    if close_idx is None:
        return s
    return s[:m.start()] + s[open_idx + 1:close_idx] + s[close_idx + 1:]


def normalize_latex_number(s: str) -> str:
    """Strip LaTeX boilerplate that doesn't change value."""
    s = s.strip()
    # Remove \boxed{...} (brace-balanced, so \boxed{\frac{1}{2}} survives)
    s = strip_boxed(s)
    # Remove dollar signs
    s = s.replace('$', '').strip()
    # Normalize whitespace
    s = re.sub(r'\s+', ' ', s)
    return s


def try_numeric(s: str) -> Optional[float]:
    """
    Try to evaluate s as a number.
    Returns float or None if not parseable as a pure number.
    """
    s = normalize_latex_number(s)
    # Handle percentages: "50%" → 0.5
    if s.endswith('%'):
        try:
            return float(s[:-1]) / 100.0
        except ValueError:
            pass
    # Handle fractions: "3/4"
    try:
        return float(Fraction(s))
    except (ValueError, ZeroDivisionError):
        pass
    # Plain float
    try:
        return float(s)
    except ValueError:
        pass
    return None


def try_sympy(s: str) -> Optional[sympy.Expr]:
    """Parse a LaTeX string with sympy."""
    try:
        expr = parse_latex(normalize_latex_number(s))
        return expr
    except Exception:
        return None


def math_equivalent(pred: str, gold: str, rtol: float = 1e-6) -> bool:
    """
    Return True iff pred and gold represent the same mathematical value.

    Strategy:
      1. Exact string match after normalization
      2. Both are numeric → compare with relative tolerance
      3. Both parse as sympy expressions → simplify(pred - gold) == 0
      4. Sympy + numeric fallback (N(expr))
    """
    pred = normalize_latex_number(pred)
    gold = normalize_latex_number(gold)

    # 1. Exact match
    if pred == gold:
        return True

    # 2. Numeric
    pn = try_numeric(pred)
    gn = try_numeric(gold)
    if pn is not None and gn is not None:
        if gn == 0:
            return abs(pn) < 1e-9
        return abs(pn - gn) / (abs(gn) + 1e-12) < rtol

    # 3. Symbolic
    pe = try_sympy(pred)
    ge = try_sympy(gold)
    if pe is not None and ge is not None:
        diff = simplify(pe - ge)
        if diff == 0:
            return True
        # Numeric fallback: evaluate to float
        try:
            pf = float(N(pe, 30))
            gf = float(N(ge, 30))
            if gf == 0:
                return abs(pf) < 1e-9
            return abs(pf - gf) / (abs(gf) + 1e-12) < rtol
        except Exception:
            pass

    return False


def extract_answer_from_completion(completion: str) -> str:
    """
    Extract the model's final answer.
    Expects the model to write \\boxed{answer} or <answer>...</answer>.
    Returns empty string if not found.
    """
    # Try \boxed{...} (LaTeX math). Search the LAST \boxed{}: models often
    # restate the format in their reasoning before committing to an answer.
    starts = [m.start() for m in re.finditer(r'\\boxed\s*\{', completion)]
    for start in reversed(starts):
        open_idx = completion.index('{', start)
        close_idx = match_braces(completion, open_idx)
        if close_idx is not None:
            return completion[open_idx + 1:close_idx].strip()
    # Try <answer>...</answer> tag
    m = re.search(r'<answer>(.*?)</answer>', completion, re.DOTALL)
    if m:
        return m.group(1).strip()
    # Fallback: last number on a "The answer is X" line
    m = re.search(r'[Tt]he answer is[:\s]+([^\n.]+)', completion)
    if m:
        return m.group(1).strip()
    return ""


def math_reward(prompt: str, completion: str, gold_answer: str) -> float:
    """
    Main entry point: return reward in {0.0, 0.1, 1.0} for a math response.
    A small partial reward of 0.1 is returned when the format is correct
    but the answer is wrong, to encourage the model to use \boxed{}.
    """
    pred = extract_answer_from_completion(completion)
    if not pred:
        return 0.0  # No answer extracted

    if math_equivalent(pred, gold_answer):
        return 1.0
    else:
        # Small reward for correct format, zero for content
        return 0.1  # format reward (optional — see section on reward shaping)

Floating-point equality traps

Never use pred == gold on raw strings. "0.333" and "1/3" are mathematically the same. "3.0" and "3" differ as strings. Always normalize before comparing, and use a tolerance for floats (relative tolerance around \(10^{-6}\) works for competition math).

Use the Library: math-verify

Writing the normalizer above is the right way to understand the problem, and you should be able to write it from scratch. In production, use math-verify (pip install math-verify) — HuggingFace’s extraction-plus-equivalence library, the grader used by lighteval and by the Open-R1 reproduction of DeepSeek-R1’s math rewards. It supersedes the hand-copied hendrycks_math normalizers that most eval harnesses used to vendor.

# pip install math-verify
from math_verify import parse, verify

gold = parse(r"$\frac{1}{2}$")                      # list of candidate readings
pred = parse(r"So the answer is $\boxed{0.5}$.")    # extraction + parsing in one call
print(verify(gold, pred))                           # True: 1/2 == 0.5

parse() performs the extraction step (\boxed{}, $...$, “the answer is …”) and deliberately returns several candidate interpretations of the same string; verify(gold, pred) reports whether any predicted candidate matches any gold candidate, using sympy for symbolic forms and set/interval-aware comparison for answers like {1,2} or (0,\infty). Because it is a maintained library, its edge-case fixes accrue to you instead of rotting in your repo.

Silent verifier death: the sympy LaTeX backend

sympy.parsing.latex.parse_latex needs an optional parser backend — the antlr4-python3-runtime package (or backend="lark" with lark installed). If it is missing, try_sympy raises inside its except Exception and returns None for every input, so the symbolic branch never fires and you are silently running a numeric-only verifier. Assert on a known symbolic pair (e.g. math_equivalent(r"\sqrt{4}", "2")) in a unit test so this fails loudly at CI time rather than as a mysterious accuracy plateau at step 400.

How Good Is Your Verifier? Measuring False Negatives

A verifier is a classifier, so it has a precision/recall profile, and you should measure it before spending GPU-hours on it. Build a small audit set — 200–500 (gold, model-answer) pairs sampled from real rollouts and labelled by hand — and report two numbers: the false-negative rate (correct answers scored 0, usually a formatting or parser gap) and the false-positive rate (wrong answers scored 1, usually an over-eager substring match).

False negatives are the more damaging failure in group-relative RL. Consider a GRPO group of \(G = 8\) completions where 4 are genuinely correct. With a perfect verifier the rewards are four 1s and four 0s, \(\mu = 0.5\), and every correct rollout gets a positive advantage. Now let the verifier reject one correct answer: rewards become three 1s and five 0s, \(\mu = 0.375\), and that rejected-but-correct rollout receives advantage \((0 - 0.375)/\sigma \approx -0.77\). The gradient does not merely ignore a good solution — it actively pushes the policy away from it, and away from whatever answer format triggered the parse failure. Huang et al. (2025), cited in the SoTA box, show this rate is not static: rule-based false negatives grow as the policy improves and starts emitting more exotic-but-valid forms, which is why verifier auditing is a recurring task, not a one-time setup step.

Cheap continuous audit

Log, for every training step, the fraction of completions where the answer extractor returned an empty string but the completion is long and terminated normally. That statistic needs no labels, costs nothing, and spikes exactly when the policy drifts into a format your extractor does not parse.

Code Execution Verifiers

For coding tasks, the verifier runs unit tests inside a sandbox and returns pass/fail. The reward can be binary (all tests pass = 1.0) or proportional (fraction of tests passed).

"""
code_verifier.py — unit-test-based reward for code generation.

This module orchestrates sandboxed execution (see sandboxes section)
and returns a float reward based on test outcomes.
"""

import re
import subprocess
import sys
import tempfile
import os
import json
import textwrap
from dataclasses import dataclass


@dataclass
class TestResult:
    passed: int         # number of tests that passed
    total: int          # total tests
    timed_out: bool     # did any test time out?
    error: str          # compilation/import error if any


def run_tests_in_subprocess(
    code: str,
    test_suite: str,
    timeout_seconds: float = 5.0,
) -> TestResult:
    """
    Execute generated code + test suite in a temporary directory.
    Returns TestResult.

    IMPORTANT: For production use, wrap this in a proper sandbox
    (see the Sandboxed Execution section). This bare subprocess
    approach is only safe on air-gapped machines or inside a container.
    """
    with tempfile.TemporaryDirectory() as tmpdir:
        # Write generated code
        code_path = os.path.join(tmpdir, "solution.py")
        with open(code_path, "w") as f:
            f.write(code)

        # Write test runner that imports solution.py.
        # NOTE: build the template flush-left and indent the embedded
        # test_suite block explicitly — mixing textwrap.dedent() with a
        # multi-line f-string substitution is fragile: dedent() computes a
        # single common leading-whitespace prefix over the *whole* string,
        # but only the first line of an embedded multi-line value inherits
        # the template's indentation (subsequent lines don't), so the
        # inserted test functions end up misaligned relative to the
        # surrounding `try:` block and raise IndentationError.
        indented_tests = textwrap.indent(test_suite, "    ")
        runner = f'''import sys, json, traceback
sys.path.insert(0, {repr(tmpdir)})

results = {{"passed": 0, "total": 0, "error": ""}}
try:
    from solution import *
{indented_tests}
    # Each test function is test_<name>; discover and run
    test_fns = [v for k, v in globals().items() if k.startswith("test_")]
    results["total"] = len(test_fns)
    for fn in test_fns:
        try:
            fn()
            results["passed"] += 1
        except Exception:
            pass
except Exception as e:
    results["error"] = traceback.format_exc()
    results["total"] = 1  # At least one test failed

print(json.dumps(results))
'''
        runner_path = os.path.join(tmpdir, "runner.py")
        with open(runner_path, "w") as f:
            f.write(runner)

        try:
            proc = subprocess.run(
                # Use sys.executable, not a bare "python": many Linux/CI
                # images only have "python3" on PATH, so a hardcoded
                # "python" raises FileNotFoundError there.
                [sys.executable, runner_path],
                capture_output=True, text=True,
                timeout=timeout_seconds,
            )
            if proc.returncode != 0 and not proc.stdout:
                return TestResult(0, 1, False, proc.stderr[:500])
            data = json.loads(proc.stdout.strip())
            return TestResult(
                passed=data["passed"],
                total=max(data["total"], 1),
                timed_out=False,
                error=data.get("error", ""),
            )
        except subprocess.TimeoutExpired:
            return TestResult(0, 1, timed_out=True, error="timeout")
        except json.JSONDecodeError as e:
            return TestResult(0, 1, False, f"JSON decode error: {e}")


def extract_code_block(completion: str) -> str:
    """
    Pull the code out of a chat completion. Models wrap solutions in markdown
    fences; take the LAST fenced block, since a model often shows a failed
    attempt first and the final block is the one it stands behind.
    Falls back to the raw completion if there is no fence.
    """
    # Build the fence pattern from a variable rather than writing three
    # literal backticks: this file is itself embedded in markdown, and a
    # literal fence inside a fence terminates the block early.
    fence = "`" * 3
    pattern = fence + r"(?:python|py)?\n(.*?)" + fence
    blocks = re.findall(pattern, completion, re.DOTALL)
    return blocks[-1] if blocks else completion


def code_reward(
    prompt: str,
    completion: str,
    test_suite: str,
    partial_credit: bool = True,
) -> float:
    """
    Reward for code generation:
      1.0  if all tests pass
      k/n  if partial_credit and k of n tests pass
      0.0  if syntax error, timeout, or all tests fail
    """
    code = extract_code_block(completion)

    result = run_tests_in_subprocess(code, test_suite)

    if result.timed_out:
        return 0.0
    if result.error and result.passed == 0:
        return 0.0
    if partial_credit:
        return result.passed / result.total
    else:
        return 1.0 if result.passed == result.total else 0.0

Sandboxed Execution: Security, Isolation & Throughput

Running arbitrary model-generated code is a serious security risk. A model that has learned to manipulate the reward function will eventually generate code that:

  • Reads its own test file and hard-codes the expected output.
  • Calls os.kill(os.getpid(), signal.SIGTERM) to time out, then retries with correct answers.
  • Escapes to the host via container breakouts.
  • Consumes unbounded CPU/memory, starving other workers.

A production sandbox must enforce: no network access, no filesystem writes outside tmpdir, CPU and memory limits, syscall filtering, and process tree isolation.

Sandbox Architectures

Option A — Docker per-sample (simple, slow) Reward Server (Python) docker run --rm ... generated code + tests ~200-500 ms per sample Isolation: strong (full container) | Security: good Simplest to set up; per-run container startup cost dominates at scale. PREFERRED Option B — Reusable microVM (fast, preferred for scale) Reward Server gRPC Firecracker MicroVM Pool — snapshotted state ~10-30 ms per sample Isolation: strong (hardware VM) | Security: very good | Speed: fast Snapshot restore reuses clean state; warm pool eliminates per-run startup cost. Size pool = throughput / per-worker capacity, with 2-3x headroom. Option C — seccomp + cgroups in-process   (fast, riskier) Reward Worker Process fork() -> install seccomp-BPF -> exec code cgroup v2: 512 MB RAM, 1 CPU, 5s wall-clock fastest (in-process) Isolation: partial (same OS kernel) | Security: moderate (kernel exploits possible) Trusted infra only; seccomp-BPF filters syscalls; cgroup v2 caps CPU/RAM/wall-time. isolation strength: A (strongest) -- B -- C (weakest) speed: A (slowest) -- B -- C (fastest)
Three sandbox architectures for code-execution rewards, ordered by isolation vs. speed tradeoff. Option A (Docker per-sample) is the simplest but incurs 200-500 ms per-run overhead. Option B (Firecracker microVM warm pool, gRPC) is the recommended production choice at 10-30 ms per sample. Option C (seccomp+cgroups in-process) is fastest but provides only partial kernel-level isolation.

For training at scale (on the order of thousands of completions per second), Docker-per-sample is too slow. The standard approach is a warm pool of microVMs or containers that are snapshotted after initialization, then restored to a clean state between runs.

You do not have to build the isolation layer yourself; each rung of the ladder has a real, well-maintained implementation:

Isolation layer Open-source tool Cost per run What it stops
In-process guards openai/human-eval’s reliability_guard in execution.py ~0 Accidental os.remove, os.kill, shutil.rmtree; not a security boundary
Namespaces + seccomp google/nsjail, containers/bubblewrap (bwrap) ~1–10 ms Filesystem, network, PID escapes; syscall surface
Container Docker / Podman with --network=none --read-only --pids-limit ~100–300 ms cold Everything above, plus resource limits via cgroups
Userspace kernel google/gvisor (docker run --runtime=runsc …) ~100–300 ms cold Kernel-exploit escapes: syscalls hit gVisor’s Go kernel, not the host’s
microVM firecracker-microvm/firecracker, and E2B which wraps it behind an SDK ~150 ms boot, snapshot-restore faster Hardware-virtualization boundary; strongest practical isolation

The pragmatic 2026 default for an RLVR run is gVisor or Firecracker underneath a warm pool: keep the pool code below exactly as written and change only the docker run line (--runtime=runsc) or swap the subprocess for an E2B Sandbox handle. Do not skip the layer entirely — the reliability_guard approach is deliberately shipped commented-out in human-eval precisely because monkey-patching os inside the same interpreter is trivially reversible by generated code.

Building a Sandboxed Execution Service

"""
sandbox_pool.py — reusable sandbox pool using Docker containers.

In production you would use Firecracker or gVisor microVMs;
Docker containers are shown here for clarity.

Each sandbox is a long-lived container. We send code + tests
over stdin and get results back over stdout, avoiding the per-run
container startup cost (~200ms) by reusing warm containers.
"""

import threading
import queue
import select
import subprocess
import json
import time
from dataclasses import dataclass, field
from typing import Optional


SANDBOX_IMAGE = "python:3.11-slim"
# Dockerfile for the sandbox image would add no extra packages
# and run as a non-root user, but we keep it simple here.

SANDBOX_RUNNER = """
import sys, json, traceback, signal, resource

# Hard memory limit: 256 MB
resource.setrlimit(resource.RLIMIT_AS, (256 * 1024 * 1024, 256 * 1024 * 1024))

# Read tasks from stdin until EOF
for line in sys.stdin:
    task = json.loads(line)
    code = task["code"]
    tests = task["tests"]
    result = {"passed": 0, "total": 0, "error": ""}

    try:
        exec_globals = {}
        exec(compile(code, "<generated>", "exec"), exec_globals)
        exec(compile(tests, "<tests>", "exec"), exec_globals)
        fns = [v for k, v in exec_globals.items() if k.startswith("test_")]
        result["total"] = len(fns)
        for fn in fns:
            try:
                fn()
                result["passed"] += 1
            except Exception:
                pass
    except Exception:
        result["error"] = traceback.format_exc()[-500:]
        result["total"] = 1

    print(json.dumps(result), flush=True)
"""


@dataclass
class SandboxWorker:
    proc: subprocess.Popen
    lock: threading.Lock = field(default_factory=threading.Lock)
    last_used: float = field(default_factory=time.time)


class SandboxPool:
    """
    A pool of warm sandbox processes.
    Thread-safe: acquire() blocks until a worker is free.
    """

    def __init__(self, pool_size: int = 8, timeout: float = 5.0):
        self.timeout = timeout
        self._available: queue.Queue = queue.Queue()
        self._all_workers = []
        for _ in range(pool_size):
            w = self._spawn()
            self._all_workers.append(w)
            self._available.put(w)

    def _spawn(self) -> SandboxWorker:
        """Start a sandbox process that reads JSON tasks from stdin."""
        proc = subprocess.Popen(
            [
                "docker", "run", "--rm", "--interactive",
                "--network=none",           # No network
                "--memory=256m",            # 256 MB RAM limit
                "--cpus=1",                 # 1 vCPU
                "--pids-limit=50",          # No fork bombs
                "--read-only",              # Read-only root filesystem
                "--tmpfs=/tmp:size=64m",    # Small writable /tmp
                SANDBOX_IMAGE,
                "python", "-c", SANDBOX_RUNNER,
            ],
            stdin=subprocess.PIPE,
            stdout=subprocess.PIPE,
            stderr=subprocess.DEVNULL,
            text=True,
        )
        return SandboxWorker(proc=proc)

    def execute(self, code: str, tests: str) -> dict:
        """
        Run code + tests in a sandbox. Blocks until a worker is available.
        Returns {"passed": int, "total": int, "error": str}.
        """
        worker: SandboxWorker = self._available.get(timeout=30.0)
        try:
            task = json.dumps({"code": code, "tests": tests}) + "\n"
            worker.proc.stdin.write(task)
            worker.proc.stdin.flush()
            # Read response with a timeout. A subprocess.Popen(text=True)
            # pipe is a TextIOWrapper over a plain OS pipe, not a socket,
            # so it has no `_sock` attribute to call settimeout() on —
            # select() is the correct way to bound the wait on a pipe fd.
            ready, _, _ = select.select([worker.proc.stdout], [], [], self.timeout)
            if not ready:
                raise TimeoutError(f"sandbox worker timed out after {self.timeout}s")
            line = worker.proc.stdout.readline()
            return json.loads(line)
        except Exception as e:
            # Worker is dead (or wedged); kill and replace it, keeping the
            # registry in sync so shutdown can reap every live container.
            worker.proc.kill()
            self._all_workers.remove(worker)
            worker = self._spawn()
            self._all_workers.append(worker)
            return {"passed": 0, "total": 1, "error": str(e)}
        finally:
            worker.last_used = time.time()
            self._available.put(worker)  # Return worker to pool

Practical sandbox tuning

For RLVR training at scale (e.g., 4096 rollout completions per training step), you typically need a pool of 64-256 sandbox workers per GPU node, each handling ~20 requests/second. The bottleneck shifts from process startup (eliminated by warm pools) to test suite I/O. Keep test suites small (under 1 KB) and use in-memory temp files rather than disk I/O.

Throughput Arithmetic

Sandbox throughput sizing

Suppose your training run generates 4096 completions per step, each requiring ~50 ms of sandbox execution (10 unit tests, each 5 ms), and a new step begins every 30 seconds.

Required throughput: \(4096 / 30 \approx 137\) completions/second.

Per worker capacity: \(1000 \text{ ms} / 50 \text{ ms} = 20\) completions/second.

Workers needed: \(\lceil 137 / 20 \rceil = 7\) workers.

In practice you want \(2\text{–}3\times\) headroom for variance and retries, so ~20 sandbox workers suffices for this configuration. At 256 MB RAM each, that is 5 GB of sandbox overhead on the reward server — negligible compared to the GPU memory.

If test suites are more expensive (e.g., compiling C++ or running integration tests at 2 seconds each), the same arithmetic gives ~820 workers, which requires a dedicated cluster of CPU machines, exactly the architecture used by competitive coding RL systems.

LLM-Judge Rewards

For open-ended tasks where there is no crisp ground truth — writing quality, helpfulness, reasoning coherence — we need a second LLM to score completions. This is sometimes called RLAIF (RL from AI Feedback), covered in depth in Constitutional AI, RLAIF & Self-Improvement.

Designing a Judge Prompt

A good judge prompt: 1. Specifies the evaluation criteria precisely (correctness, clarity, step-by-step validity). 2. Uses a reference answer or rubric when available. 3. Asks the judge to reason before scoring (chain-of-thought in the judge improves calibration). 4. Outputs a structured response to make score extraction reliable.

"""
llm_judge.py — LLM-as-a-judge reward for open-ended tasks.

Speaks the OpenAI chat-completions API, which is also what vLLM, SGLang and
TGI expose. Default here is a *local* judge: `vllm serve <model> --port 8000`
gives you an OpenAI-compatible /v1 endpoint, so an RL run can call a judge
thousands of times per step at zero marginal cost and with no rate limits.
Point JUDGE_BASE_URL at a hosted provider instead if you want a frontier judge.
"""

import os
import re
import asyncio
import openai
from typing import Optional


# One client per process. Constructing an AsyncOpenAI inside the request path
# creates (and leaks) a fresh connection pool on every judged completion.
JUDGE_BASE_URL = os.environ.get("JUDGE_BASE_URL", "http://localhost:8000/v1")
JUDGE_MODEL = os.environ.get("JUDGE_MODEL", "local-judge")  # --served-model-name
_client = openai.AsyncOpenAI(
    base_url=JUDGE_BASE_URL,
    api_key=os.environ.get("OPENAI_API_KEY", "EMPTY"),  # vLLM ignores the key
)


JUDGE_SYSTEM = """
You are an expert evaluator for mathematical reasoning and problem solving.
You will be given a problem, a reference solution, and a model response.
Your task is to evaluate the model response on a scale of 0 to 10.

Evaluation criteria:
- Correctness (5 pts): Is the final answer correct?
- Reasoning quality (3 pts): Is the chain-of-thought valid, with each step justified?
- Clarity (2 pts): Is the solution well-organized and readable?

Output format:
<thinking>
[Your analysis of the response]
</thinking>
<score>N</score>

Where N is an integer from 0 to 10.
""".strip()


JUDGE_USER_TEMPLATE = """
PROBLEM:
{problem}

REFERENCE SOLUTION:
{reference}

MODEL RESPONSE:
{response}

Please evaluate the model response using the criteria above.
""".strip()


async def judge_single(
    problem: str,
    response: str,
    reference: str,
    judge_model: str = JUDGE_MODEL,
    temperature: float = 0.0,
) -> float:
    """
    Returns a normalized reward in [0, 1] for a single (problem, response) pair.
    Calls the judge LLM asynchronously.
    """
    messages = [
        {"role": "system", "content": JUDGE_SYSTEM},
        {"role": "user", "content": JUDGE_USER_TEMPLATE.format(
            problem=problem,
            reference=reference,
            response=response,
        )},
    ]

    resp = await _client.chat.completions.create(
        model=judge_model,
        messages=messages,
        temperature=temperature,
        max_tokens=512,
    )

    text = resp.choices[0].message.content
    m = re.search(r"<score>(\d+(?:\.\d+)?)</score>", text)
    if m:
        raw_score = float(m.group(1))
        return min(max(raw_score / 10.0, 0.0), 1.0)

    # Fallback: try to parse any integer at the end of the response
    nums = re.findall(r"\b(\d+)\b", text)
    if nums:
        raw = float(nums[-1])
        return min(max(raw / 10.0, 0.0), 1.0)

    return 0.5  # Uncertain — return neutral score


async def batch_judge(
    problems: list[str],
    responses: list[str],
    references: list[str],
    judge_model: str = JUDGE_MODEL,
    concurrency: int = 32,
) -> list[float]:
    """
    Evaluate a batch of completions concurrently.
    concurrency limits simultaneous API calls to avoid rate limits.
    """
    sem = asyncio.Semaphore(concurrency)

    async def bounded_judge(p, r, ref):
        async with sem:
            return await judge_single(p, r, ref, judge_model)

    tasks = [
        bounded_judge(p, r, ref)
        for p, r, ref in zip(problems, responses, references)
    ]
    return await asyncio.gather(*tasks)

Judge Calibration and Positional Bias

LLM judges suffer from positional bias (they prefer whichever answer appears first), verbosity bias (longer responses score higher regardless of quality), and self-preference bias (a judge from the same family as the policy gives spuriously high scores). Mitigation strategies:

  • Swap the order of responses in pairwise comparisons and average the scores.
  • Use a judge from a different model family than the policy.
  • Normalize scores within each prompt batch: \(\tilde{r}_i = (r_i - \mu_{\text{batch}}) / \sigma_{\text{batch}}\).
  • Include a “null” baseline completion (e.g., “I don’t know”) to anchor the scale.

Reward Shaping and the Format vs. Correctness Tradeoff

Raw binary rewards (correct/incorrect) create sparse gradients, especially early in training when the model rarely reaches the right answer. Reward shaping introduces auxiliary rewards to guide the model toward behaviors that are likely to lead to the correct answer.

Format Rewards

A common technique in RLVR is to give a small positive reward for using the expected output format (e.g., a <think>...</think> reasoning block followed by a \boxed{answer}), even when the answer is wrong. This solves a cold-start problem: the model must first learn to produce structured output before it can receive meaningful correctness rewards.

\[ r_{\text{total}} = r_{\text{correctness}} + \lambda_f \cdot r_{\text{format}} \]
r_total = r_correctness + lambda_f * r_format Healthy shaping (small lambda_f ~= 0.1) training steps reward format reward (cold-start solved) correctness (lifts off after format, tracks it) <think> ...valid reasoning... </think> \boxed{correct} format and correctness rise together Gaming (lambda_f too large) training steps gap: decoupled format reward (pegs ~= 100%) correctness (stuck ~low) <think> ...anything... </think> \boxed{0} perfect format, wrong answer, still rewarded Rule of thumb Track corr(format_reward, correctness_reward) throughout training. If the two curves decouple (right panel), reduce lambda_f. Typical healthy range: lambda_f ~= 0.1 - 0.2.
A small format weight solves the cold-start problem; too large a one lets the model game format instead of answering correctly. With lambda_f around 0.1, format reward saturates first and correctness reward follows and keeps climbing with it; with lambda_f too large, format reward pegs near its ceiling while correctness stalls, and the model converges on a degenerate template like \boxed{0} that always collects the format bonus — the fix is to monitor the correlation between the two reward streams and shrink lambda_f if they decouple.

where \(r_{\text{format}} \in \{0, 1\}\) indicates whether the output satisfies the format, and \(\lambda_f\) is typically small (on the order of 0.1 to 0.2) to avoid the model gaming format at the expense of correctness.

Format reward gaming

If \(\lambda_f\) is too large, the model learns to produce syntactically correct but semantically empty responses. For example, it may always write <think>Let me solve this step by step.</think>\boxed{0} — perfect format, always wrong answer, but nonzero reward. Monitor the correlation between format reward and correctness reward throughout training; if they decouple, reduce \(\lambda_f\).

Process Reward Models (PRMs)

Instead of rewarding only the final answer, a Process Reward Model (PRM) assigns a reward to each step in the chain-of-thought. This is covered in depth in RL with Verifiable Rewards (RLVR) & The Reasoning Recipe; here we focus on the infrastructure interface.

"""
prm_reward.py — step-level reward using a process reward model.

The PRM is a separate model (often a smaller classifier fine-tuned
on step-level correctness annotations) that scores each reasoning step.
"""

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification


class ProcessRewardModel:
    """
    Wraps a PRM that scores individual reasoning steps.
    Input: a (problem, partial_solution_so_far) pair.
    Output: a scalar score in [0, 1] for the most recent step.
    """

    def __init__(self, model_name: str, device: str = "cuda"):
        self.tokenizer = AutoTokenizer.from_pretrained(model_name)
        self.model = AutoModelForSequenceClassification.from_pretrained(
            model_name, num_labels=1,
        ).to(device).eval()
        self.device = device

    @torch.no_grad()
    def score_steps(
        self,
        problem: str,
        steps: list[str],
        batch_size: int = 32,
    ) -> list[float]:
        """
        Score each step in a chain-of-thought.
        steps: list of individual reasoning steps (not cumulative).
        Returns a list of scalar rewards, one per step.
        """
        scores = []
        # Build cumulative context for each step
        contexts = []
        for i in range(len(steps)):
            # Include all steps up to and including step i
            partial = problem + "\n\n" + "\n".join(steps[:i+1])
            contexts.append(partial)

        # Batch inference
        for start in range(0, len(contexts), batch_size):
            batch = contexts[start:start + batch_size]
            enc = self.tokenizer(
                batch,
                return_tensors="pt",
                padding=True,
                truncation=True,
                max_length=2048,
            ).to(self.device)
            logits = self.model(**enc).logits.squeeze(-1)
            # Sigmoid to get probability of "step is correct"
            probs = torch.sigmoid(logits).cpu().tolist()
            if isinstance(probs, float):
                probs = [probs]
            scores.extend(probs)

        return scores

    def aggregate(
        self,
        step_scores: list[float],
        method: str = "min",
    ) -> float:
        """
        Aggregate step scores into a single episode reward.

        'min': reward = min step score (weakest-link; used in Math-Shepherd)
        'last': reward = score of final step (common in practice)
        'mean': reward = mean of all step scores
        """
        if not step_scores:
            return 0.0
        if method == "min":
            return min(step_scores)
        elif method == "last":
            return step_scores[-1]
        elif method == "mean":
            return sum(step_scores) / len(step_scores)
        raise ValueError(f"Unknown aggregation method: {method}")

Multi-Objective and Weighted Rewards

Most production RLVR runs combine multiple reward signals:

Signal Weight Purpose
Correctness (rule-based) 1.0 Primary learning signal
Format compliance 0.1–0.2 Cold-start / structure
Reasoning length penalty -0.01 per token over limit Control verbosity
Safety classifier -5.0 if flagged Hard constraint
LLM judge (helpfulness) 0.3 Open-ended quality

The weighted combination is:

\[ r_{\text{total}} = \sum_{k} w_k \cdot r_k \]

This is simple but has pitfalls. If reward magnitudes differ by orders of magnitude (e.g., correctness ∈ {0, 1} but a continuous fluency score ∈ [0, 100]), one term will dominate the gradient. Always normalize rewards to similar scales before combining.

Reward Normalization Per Prompt

A robust pattern used in practice is to normalize rewards within the batch of completions for a single prompt (a “prompt group”):

\[ \tilde{r}_{i} = \frac{r_i - \mu_{\text{group}}}{\sigma_{\text{group}} + \epsilon} \]

This is exactly the normalization used in GRPO (Group Relative Policy Optimization), discussed in GRPO, RLOO & Critic-Free RL. It ensures that the advantage signal is zero-mean per prompt regardless of the absolute reward scale.

"""
reward_combiner.py — multi-objective reward combining and normalization.
"""

import numpy as np
from dataclasses import dataclass
from typing import Callable


@dataclass
class RewardComponent:
    name: str
    fn: Callable                  # (prompt, completion, metadata) -> float
    weight: float                 # Relative importance
    clip_min: float = -10.0       # Clip extreme values
    clip_max: float = 10.0


class MultiObjectiveReward:
    """
    Combines multiple reward signals, normalizes per group,
    and returns the final scalar reward for each completion.
    """

    def __init__(self, components: list[RewardComponent]):
        self.components = components

    def compute(
        self,
        prompts: list[str],
        completions: list[str],
        metadata: list[dict],
        normalize_per_group: bool = True,
    ) -> tuple[list[float], dict[str, list[float]]]:
        """
        Returns (total_rewards, per_component_rewards).

        Assumes each (prompt, completions) group has the same prompt
        and completions are ordered: all completions for prompt 0,
        then all for prompt 1, etc.

        metadata: list of dicts with per-sample info (e.g., gold_answer).
        normalize_per_group: apply Z-score normalization within groups.
        """
        n = len(completions)
        component_rewards = {c.name: np.zeros(n) for c in self.components}

        # Evaluate each component
        for comp in self.components:
            for i, (prompt, completion, meta) in enumerate(
                zip(prompts, completions, metadata)
            ):
                raw = comp.fn(prompt, completion, meta)
                clipped = np.clip(raw, comp.clip_min, comp.clip_max)
                component_rewards[comp.name][i] = clipped

        # Weighted sum
        total = np.zeros(n)
        for comp in self.components:
            total += comp.weight * component_rewards[comp.name]

        # Normalize per group if requested
        if normalize_per_group:
            # Group by unique prompt (assumes they come in contiguous blocks)
            unique_prompts = []
            seen = {}
            for p in prompts:
                if p not in seen:
                    seen[p] = len(unique_prompts)
                    unique_prompts.append(p)
            for prompt in unique_prompts:
                idx = [i for i, p in enumerate(prompts) if p == prompt]
                group = total[idx]
                mu, sigma = group.mean(), group.std()
                total[idx] = (group - mu) / (sigma + 1e-8)

        return total.tolist(), {k: v.tolist() for k, v in component_rewards.items()}

Interview Corner

Q: In an RLVR training run for math reasoning, the model’s correctness reward plateaus at 30% after 500 steps, but format reward is 100%. What are your next steps?

A: Several failure modes are consistent with this pattern:

  1. The format reward is too large relative to correctness. If \(\lambda_f = 0.5\) and correctness rewards are sparse (0 or 1), the model may have found a local optimum where format reward alone gives a good expected return. Try reducing \(\lambda_f\) to 0.05-0.1 and restarting.

  2. The problem distribution is too hard. If the base model never produces correct answers during rollout (the oracle pass rate is near zero), there is no positive reward signal to latch onto. Solutions: use a warmer sampling temperature (e.g., 0.9 instead of 0.6) to increase exploration, curriculum with easier problems first, or start from a stronger SFT checkpoint.

  3. The math verifier has bugs. Check the false-negative rate of your verifier on known-correct answers. A buggy normalizer that rejects \frac{1}{2} when the gold is 0.5 creates a phantom ceiling. Log verifier inputs/outputs during training.

  4. KL penalty is too strong. If the KL coefficient against the reference policy is large, the policy cannot move far enough from the SFT model to find new correct solutions. Check the KL term in the loss; see Advantage Estimation, KL Control & Stability Tricks.

Plugging a Reward Function Into Real Trainers

None of the code above is useful until a trainer calls it. The two libraries you are most likely to use — TRL and veRL, covered in TRL: HuggingFace’s RL Library and veRL: HybridFlow & The Single-Controller Architecture — expose almost the same contract, and knowing it is the difference between “I have a verifier” and “I have a training run.”

TRL. GRPOTrainer takes a list of plain Python callables in reward_funcs. Each is called once per generation batch with keyword arguments prompts, completions, and every extra column of your dataset (so a gold_answer column arrives as a gold_answer= kwarg), and must return one float per completion. Multiple functions are combined with GRPOConfig(reward_weights=[...]) — that is where the multi-objective weighting of the previous section actually lives, and TRL logs each component separately so you can watch format and correctness decouple.

from datasets import Dataset
from trl import GRPOConfig, GRPOTrainer
from math_verifier import math_reward, extract_answer_from_completion

def correctness_reward(completions, gold_answer, **kwargs) -> list[float]:
    """One float per completion. `gold_answer` comes from the dataset column."""
    texts = [c[0]["content"] if isinstance(c, list) else c for c in completions]
    return [math_reward("", t, g) for t, g in zip(texts, gold_answer)]

def format_reward(completions, **kwargs) -> list[float]:
    texts = [c[0]["content"] if isinstance(c, list) else c for c in completions]
    return [1.0 if extract_answer_from_completion(t) else 0.0 for t in texts]

ds = Dataset.from_dict({
    "prompt": ["What is 1/2 + 1/2?"],
    "gold_answer": ["1"],
})

cfg = GRPOConfig(
    output_dir="grpo-math",
    num_generations=8,              # group size G for the group-relative baseline
    reward_weights=[1.0, 0.1],      # correctness weight 1.0, format weight lambda_f
    max_completion_length=1024,
)
trainer = GRPOTrainer(
    model="HuggingFaceTB/SmolLM2-135M-Instruct",
    reward_funcs=[correctness_reward, format_reward],
    args=cfg,
    train_dataset=ds,
)
# trainer.train()

veRL. Instead of passing callables, you point config at a file: custom_reward_function.path=/path/to/reward.py and custom_reward_function.name=compute_score. veRL calls it per sample with the fields it carries through the dataset, conventionally compute_score(data_source, solution_str, ground_truth, extra_info=None), returning a float (or a dict with a "score" key plus extra metrics to log). The data_source argument is what lets one run mix math, code, and instruction-following prompts through a single dispatching verifier:

# reward.py — veRL-style per-sample scorer
from math_verifier import math_reward
from code_verifier import code_reward

def compute_score(data_source, solution_str, ground_truth, extra_info=None):
    if data_source.startswith("math"):
        return math_reward("", solution_str, ground_truth)
    if data_source.startswith("code"):
        return code_reward("", solution_str, ground_truth)  # ground_truth = tests
    raise ValueError(f"no verifier registered for data_source={data_source}")

Two practical notes. First, both trainers call your function on the driver process, synchronously, inside the training step — a 2-second C++ compile per completion will stall every GPU in the run, which is the entire motivation for the reward server in the next section. Second, if you want ready-made environments rather than bare reward functions, willccbb/verifiers packages prompts, rollout logic, and rubric-style multi-criterion scoring behind one interface, and plugs into GRPO-style trainers directly.

Reward Server Architecture

A reward server decouples reward computation from training, allowing:

  • Independent scaling of reward workers (especially important when rewards involve expensive sandbox execution).
  • Reward caching (identical (prompt, completion) pairs return cached results).
  • Async reward computation to hide latency behind the next rollout batch.
Reward Server (gRPC) Request Queue (Redis) Router (task type) LRU Cache SHA256(prompt + completion) cache hit -- return immediately (skip workers) math code judge Math Verifier Workers x8 Code Runner Sandbox Pool x64 LLM Judge async API calls x32 cheap, fast expensive; size pool async; rate-limited How a request flows 1. Request enters Redis queue; Router dequeues it and reads task_type. 2. Router hashes SHA256(prompt+completion) and looks up LRU Cache. Cache hit: stored scalar returned immediately, skipping all workers. 3. Cache miss: request dispatched to matching pool (math / code / judge). 4. Code Runner pool (x64) is the bottleneck; size = rollout_rate / worker_rps. 5. LLM Judge calls are async, hiding API latency behind the next rollout batch. 6. Worker results written to cache on return; training loop receives the scalar.
Internal architecture of the gRPC reward server showing request routing, caching, and parallel worker pools. A Redis queue decouples ingress from computation; the Router checks an LRU cache (SHA256-keyed) before dispatching — a cache hit returns immediately and skips all workers. On a miss, requests route to one of three sized pools; the Code Runner sandbox pool (x64) is the throughput-critical path and must be sized to match rollout rate.

Here is a minimal FastAPI reward server:

"""
reward_server.py — minimal HTTP reward server for RL training.

Deploy behind a load balancer; run multiple instances for throughput.
"""

import asyncio
import hashlib
import json
from contextlib import asynccontextmanager
from typing import Optional

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel

from math_verifier import math_reward
from code_verifier import extract_code_block
from sandbox_pool import SandboxPool
from llm_judge import batch_judge

# Global sandbox pool — initialized once at startup
_sandbox_pool: Optional[SandboxPool] = None
# Simple in-memory cache (use Redis in production)
_cache: dict[str, float] = {}


@asynccontextmanager
async def lifespan(app: FastAPI):
    """
    Startup/shutdown hook. Use lifespan, not @app.on_event("startup"):
    the on_event decorators have been deprecated since FastAPI 0.93.
    Warming the pool here means the first training step does not pay
    container boot latency.
    """
    global _sandbox_pool
    _sandbox_pool = SandboxPool(pool_size=16, timeout=10.0)
    yield
    for w in _sandbox_pool._all_workers:
        w.proc.kill()


app = FastAPI(title="Reward Server", lifespan=lifespan)


class RewardRequest(BaseModel):
    task_type: str          # "math", "code", "judge"
    prompt: str
    completion: str
    metadata: dict          # gold_answer, test_suite, reference, etc.


class RewardResponse(BaseModel):
    reward: float
    cached: bool
    breakdown: dict         # Per-component rewards for logging


def _cache_key(req: RewardRequest) -> str:
    raw = json.dumps({
        "type": req.task_type,
        "prompt": req.prompt,
        "completion": req.completion,
        "meta": req.metadata,
    }, sort_keys=True)
    return hashlib.sha256(raw.encode()).hexdigest()


@app.post("/reward", response_model=RewardResponse)
async def compute_reward(req: RewardRequest):
    key = _cache_key(req)
    if key in _cache:
        return RewardResponse(reward=_cache[key], cached=True, breakdown={})

    breakdown = {}
    total = 0.0

    if req.task_type == "math":
        gold = req.metadata.get("gold_answer", "")
        r = math_reward(req.prompt, req.completion, gold)
        breakdown["correctness"] = r
        total = r

    elif req.task_type == "code":
        tests = req.metadata.get("test_suite", "")
        code = extract_code_block(req.completion)
        # SandboxPool.execute() blocks on a pipe to a container, so hand it to
        # the default thread pool; otherwise one slow test suite freezes the
        # event loop and every other in-flight reward request with it.
        out = await asyncio.get_running_loop().run_in_executor(
            None, lambda: _sandbox_pool.execute(code, tests)
        )
        r = out["passed"] / max(out["total"], 1)
        breakdown["code_tests"] = r
        total = r

    elif req.task_type == "judge":
        reference = req.metadata.get("reference", "")
        rewards = await batch_judge(
            [req.prompt], [req.completion], [reference]
        )
        r = rewards[0]
        breakdown["llm_judge"] = r
        total = r

    else:
        raise HTTPException(400, f"Unknown task_type: {req.task_type}")

    _cache[key] = total
    return RewardResponse(reward=total, cached=False, breakdown=breakdown)


@app.post("/reward/batch")
async def compute_rewards_batch(requests: list[RewardRequest]):
    """Process a batch of reward requests concurrently."""
    tasks = [compute_reward(req) for req in requests]
    return await asyncio.gather(*tasks)

Reward Hacking and Mitigation

Reward hacking occurs when the model finds a high-reward policy that does not correspond to the intended behavior. This is covered extensively in Reward Hacking, Over-Optimization & Alignment Failures, but from an infrastructure perspective the key defenses are:

Adversarial testing of your verifier. Before training, run a fuzzer that generates edge-case strings (e.g., \boxed{}, \boxed{\infty}, extremely long LaTeX expressions, Unicode look-alikes for digits) and check that your verifier handles them correctly. A buggy verifier is itself a reward hacking surface.

Hold-out test suites. For code tasks, separate the visible test cases (used for reward) from hidden test cases (used for eval). This mirrors competitive programming practice. If the model achieves 90% on visible tests but 30% on hidden tests, it has overfit to the visible tests — reward hacking through test memorization.

Reward variance monitoring. Track \(\text{Var}(r)\) within each prompt group over training. Healthy training shows decreasing variance as the model converges. A variance spike often indicates the model discovered a reward shortcut.

KL-constrained reward. Adding a KL penalty \(-\beta \cdot D_{\text{KL}}(\pi \| \pi_{\text{ref}})\) to the reward limits how far the policy can deviate from the reference model. This is not a perfect defense, but it makes dramatic behavioral changes more expensive in reward terms. See Advantage Estimation, KL Control & Stability Tricks.

Putting It Together: A Complete Reward Pipeline

Here is how the components connect in a real RLVR system. This is precisely the pipeline the capstone runs: the narrow GRPO stage of Stack-100M in Post-Training: SFT, DPO, and Narrow RLVR (GRPO) That Works at 100M uses a correctness verifier plus a small format reward exactly as below, and the tool-call reward for its research agent — did the agent emit a well-formed call, and did the call return usable evidence — is built the same way in A Narrow Auto-Research Agent: ReAct, Tool-Use & Retrieval by Distillation. At 100M parameters the verifier’s false-negative rate matters more than at frontier scale, because the true pass rate is low enough that losing one correct rollout per group can wipe out the entire positive signal for that prompt.

"""
rl_reward_pipeline.py — end-to-end reward pipeline for a math RLVR run.

This orchestrates: rollout → reward computation → advantage normalization.
Intended to run on the CPU reward server, called from the training loop.
"""

import re
import asyncio
import numpy as np
from typing import NamedTuple


class RolloutBatch(NamedTuple):
    prompts: list[str]            # One per prompt group
    completions: list[list[str]]  # completions[i] = list of G completions for prompt i
    gold_answers: list[str]       # Reference answers for verifier


class RewardBatch(NamedTuple):
    rewards: np.ndarray           # Shape (N,) flat — all completions
    advantages: np.ndarray        # Shape (N,) after group normalization
    component_log: dict           # For logging / dashboards


async def compute_rewards_for_batch(
    batch: RolloutBatch,
    format_lambda: float = 0.1,
) -> RewardBatch:
    """
    Main pipeline:
    1. Compute correctness reward (rule-based math verifier)
    2. Compute format reward
    3. Combine with weights
    4. Normalize per group to get advantages
    """
    from math_verifier import math_reward, extract_answer_from_completion

    flat_prompts = []
    flat_completions = []
    flat_golds = []
    group_ids = []  # Which group each completion belongs to

    for g, (prompt, comps, gold) in enumerate(
        zip(batch.prompts, batch.completions, batch.gold_answers)
    ):
        for comp in comps:
            flat_prompts.append(prompt)
            flat_completions.append(comp)
            flat_golds.append(gold)
            group_ids.append(g)

    n = len(flat_completions)
    correctness = np.zeros(n)
    format_r = np.zeros(n)

    for i, (prompt, comp, gold) in enumerate(
        zip(flat_prompts, flat_completions, flat_golds)
    ):
        # Format reward: 1 if \boxed{} is present, 0 otherwise
        has_boxed = bool(re.search(r'\\boxed\s*\{', comp))
        format_r[i] = 1.0 if has_boxed else 0.0
        # Correctness reward
        correctness[i] = math_reward(prompt, comp, gold)

    combined = correctness + format_lambda * format_r

    # Normalize per group → advantages
    advantages = np.zeros(n)
    num_groups = len(batch.prompts)
    for g in range(num_groups):
        idx = [i for i, gid in enumerate(group_ids) if gid == g]
        group = combined[idx]
        mu, sigma = group.mean(), group.std()
        advantages[idx] = (group - mu) / (sigma + 1e-8)

    return RewardBatch(
        rewards=combined,
        advantages=advantages,
        component_log={
            "correctness": correctness.tolist(),
            "format": format_r.tolist(),
        },
    )

Key Takeaways

  • Rule-based verifiers (math equivalence, code unit tests) are the gold standard for RLVR: deterministic, free at inference time, and hard to hack compared to learned reward models.
  • Math verifiers must do symbolic/numeric normalization, not string equality, and must match braces rather than regex \boxed{...}. Sympy with a numeric fallback handles most competition math; in production reach for math-verify.
  • A verifier is a classifier: audit its false-negative rate on labelled rollouts — in group-relative RL a rejected-but-correct completion gets a negative advantage, actively training the policy away from a valid solution — and keep watching it with fuzzing, hold-out test suites, and within-group reward variance.
  • Trainers consume rewards through a narrow contract — TRL’s reward_funcs callables (weighted via GRPOConfig.reward_weights) and veRL’s custom_reward_function.compute_score — and they run it synchronously in the training step, which is why heavy verifiers belong behind a reward server.
  • Code verifiers need sandboxed execution to prevent reward hacking via filesystem or process manipulation. Warm container or microVM pools eliminate per-sample startup overhead.
  • LLM-judge rewards are necessary for open-ended tasks but introduce positional and verbosity bias; mitigate by swapping order, using cross-family judges, and normalizing within batches.
  • Format rewards (small \(\lambda \approx 0.1\)) solve the cold-start problem by encouraging structured output before correctness rewards become dense; too large a \(\lambda\) invites gaming.
  • Multi-objective rewards should be combined on similar scales and normalized per prompt group (GRPO-style) to produce zero-mean advantages independent of absolute reward magnitude.
  • The reward server is a critical piece of RL infrastructure: decouple it from training, cache results, and size your sandbox pool to match rollout throughput (workers = throughput / per-worker capacity, plus 2-3x headroom).

State of the Art & Resources (2026)

Reward engineering for RLVR has matured rapidly since DeepSeek-R1 demonstrated that rule-based verifiers alone—without any learned reward model—can drive state-of-the-art reasoning. The active frontier now spans more robust symbolic verifiers, step-level process reward models that scale test-time compute, and production-grade sandboxing infrastructure for safe code execution at training throughput.

Foundational work

Recent advances (2023–2026)

Open-source & tools

  • verl-project/verl — flexible, high-throughput RL post-training framework (PPO, GRPO, DAPO) with pluggable reward functions; used to reproduce and extend DeepSeek-R1.
  • huggingface/Math-Verify — the standard answer-extraction and equivalence library for math RLVR and evals (pip install math-verify); used by lighteval and Open-R1, and the maintained replacement for vendored hendrycks_math normalizers.
  • willccbb/verifiers — environment + rubric abstraction for RLVR: packages prompt, rollout, and multi-criterion scoring behind one interface that GRPO-style trainers can consume directly.
  • e2b-dev/E2B — open-source Firecracker-backed cloud sandbox SDK for safely executing AI-generated code; Python and TypeScript APIs.
  • google/gvisor — userspace kernel that gives container-speed startup with a much smaller host syscall surface; drop-in via docker run --runtime=runsc.
  • openai/prm800k — 800 K step-level correctness labels on MATH solutions plus the SymPy-based answer-grading logic from “Let’s Verify Step by Step.”
  • opendilab/awesome-RLVR — curated, actively updated reading list of RLVR papers, codebases, and tutorials (2024–2026).

Go deeper

Further Reading

  • Lightman et al., Let’s Verify Step by Step (OpenAI, 2023) — introduces process reward models (PRMs) for math reasoning and the PRM800K dataset.
  • Wang et al., Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations (2023) — automated PRM training via Monte Carlo rollouts.
  • DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025) — the canonical RLVR recipe with format + correctness rewards.
  • Chen et al., Evaluating Large Language Models Trained on Code (OpenAI, 2021) — introduces HumanEval and the pass@k metric for code reward evaluation.
  • Guo et al., Deepseek-Coder: When the Large Language Model Meets Programming (2024) — discusses code-execution reward pipelines at scale.
  • Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023) — systematic analysis of LLM judge biases and calibration.
  • Shen et al., Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback (2023) — reward shaping to control response length.
  • openai/human-eval — reference sandboxed execution harness (execution.py, reliability_guard) for code-correctness rewards; openai/evals for general eval scaffolding.
  • huggingface/Math-Verify — the maintained math answer extraction/equivalence library used by lighteval and Open-R1.

Exercises

1. The chapter’s math_reward is called on a completion ending in \boxed{0.50}, with gold_answer = "1/2". Trace math_equivalent("0.50", "1/2") through the numbered strategy in its docstring and say which branch returns True and why. Then explain in one sentence why math_reward returns 0.1 (not 0.0) for a well-formatted but wrong answer, and what training pathology the !!! warning "Format reward gaming" admonition warns this can cause if that idea is pushed too far.

Solution

First extract_answer_from_completion pulls 0.50 out of \boxed{0.50}, and math_reward calls math_equivalent("0.50", "1/2"). Both arguments are first passed through normalize_latex_number, which only strips \boxed{}, $, and whitespace, so they are unchanged.

  • Strategy 1 (exact string match): "0.50" == "1/2" is False.
  • Strategy 2 (numeric): try_numeric("0.50") is not a percentage, so it tries float(Fraction("0.50")). Fraction accepts decimal strings, giving Fraction(1, 2) -> 0.5. try_numeric("1/2") gives float(Fraction("1/2")) -> 0.5. Both are non-None, gn = 0.5 != 0, and the relative error is
\[ \frac{|0.5 - 0.5|}{|0.5| + 10^{-12}} = 0 < 10^{-6}. \]

So strategy 2 returns True; the symbolic sympy branch is never reached.

math_reward returns 0.1 on a wrong-but-formatted answer as a format reward: it gives dense signal early in training so the model first learns to emit \boxed{} before correctness rewards become reachable. Pushed too far (a large \(\lambda_f\)), the model games format at the expense of content — e.g. always emitting <think>...</think>\boxed{0} for a nonzero reward with zero correctness — which is exactly the decoupling the format-gaming warning tells you to monitor.

2. You run RLVR on a coding task. Each training step generates 8192 completions, each needs ~80 ms of sandbox execution, and a new step starts every 40 s. Using the chapter’s “Throughput arithmetic” method: (a) compute the required completions/second, (b) the per-worker capacity, © the number of workers with no headroom, and (d) a sized pool with the chapter’s recommended \(2\text{--}3\times\) headroom, plus its total RAM at 256 MB/worker. (e) If you instead compile C++ at 2 s per completion, how many workers (no headroom) does the same required throughput demand?

Solution

(a) Required throughput:

\[ 8192 / 40 = 204.8 \text{ completions/second.} \]

(b) Per-worker capacity at 80 ms each:

\[ 1000 \text{ ms} / 80 \text{ ms} = 12.5 \text{ completions/second.} \]

© Workers with no headroom:

\[ \lceil 204.8 / 12.5 \rceil = \lceil 16.4 \rceil = 17 \text{ workers.} \]

(d) With \(3\times\) headroom, \(17 \times 3 = 51\) workers (round to ~50). RAM:

\[ 51 \times 256 \text{ MB} \approx 13 \text{ GB}, \]

negligible against GPU memory, matching the chapter’s point that sandbox overhead is cheap.

(e) At 2 s/completion, per-worker capacity is \(1/2 = 0.5\) completions/second, so

\[ \lceil 204.8 / 0.5 \rceil = 410 \text{ workers}, \]

a dedicated CPU cluster — the same conclusion the chapter draws for compiled/integration-test workloads.

3. A GRPO prompt group has 4 completions. The rule-based math verifier scores correctness as [1, 0, 0, 0]; all four responses contain \boxed{}, so the format reward is [1, 1, 1, 1]. Using the pipeline’s combination \(r = r_{\text{correctness}} + \lambda_f \cdot r_{\text{format}}\) with \(\lambda_f = 0.1\), and the per-group normalization \(\tilde r_i = (r_i - \mu)/(\sigma + \epsilon)\) (use population \(\sigma\), i.e. np.std default, and ignore \(\epsilon\)), compute the four advantages by hand. What do they sum to, and why is that expected?

Solution

Combined rewards:

\[ r = [1 + 0.1,\ 0 + 0.1,\ 0 + 0.1,\ 0 + 0.1] = [1.1,\ 0.1,\ 0.1,\ 0.1]. \]

Mean:

\[ \mu = (1.1 + 0.1 + 0.1 + 0.1)/4 = 1.4/4 = 0.35. \]

Deviations: \([0.75,\ -0.25,\ -0.25,\ -0.25]\). Population variance:

\[ \sigma^2 = \frac{0.75^2 + 3\cdot(0.25)^2}{4} = \frac{0.5625 + 0.1875}{4} = \frac{0.75}{4} = 0.1875, \]

so \(\sigma = \sqrt{0.1875} \approx 0.4330\). Advantages:

\[ \tilde r = \left[\frac{0.75}{0.4330},\ \frac{-0.25}{0.4330},\ \frac{-0.25}{0.4330},\ \frac{-0.25}{0.4330}\right] \approx [1.732,\ -0.577,\ -0.577,\ -0.577]. \]

They sum to (approximately) 0: subtracting the group mean makes the advantages zero-mean by construction. That is the whole point of per-group normalization — the advantage signal is centered per prompt regardless of the absolute reward scale, so only relative quality within the group drives the update.

4. The format-gaming warning says a too-large \(\lambda_f\) lets the model win by faking format. Make it quantitative. Compare two policies under \(r = r_{\text{correctness}} + \lambda_f \cdot r_{\text{format}}\), both with \(r_{\text{correctness}}, r_{\text{format}} \in \{0,1\}\):

  • Honest: genuinely attempts problems; expected correctness \(0.2\), but its messy outputs only emit \boxed{} half the time, so expected format \(0.5\).
  • Degenerate: always emits <think>...</think>\boxed{0}; correctness \(0\), format \(1\).

Find the threshold value of \(\lambda_f\) above which the degenerate policy earns the higher expected reward. Comment on the chapter’s recommended range \(\lambda_f \in [0.1, 0.2]\).

Solution

Expected rewards:

\[ \mathbb{E}[r_{\text{honest}}] = 0.2 + \lambda_f \cdot 0.5, \qquad \mathbb{E}[r_{\text{degen}}] = 0 + \lambda_f \cdot 1. \]

Degenerate wins when

\[ \lambda_f > 0.2 + 0.5\,\lambda_f \;\Longrightarrow\; 0.5\,\lambda_f > 0.2 \;\Longrightarrow\; \lambda_f > 0.4. \]

So the gaming optimum only takes over above \(\lambda_f = 0.4\). The chapter’s recommended \(\lambda_f \in [0.1, 0.2]\) sits comfortably below this threshold: at \(\lambda_f = 0.2\), honest scores \(0.2 + 0.1 = 0.3\) versus degenerate \(0.2\), so honest still wins. This is the concrete reason the chapter keeps \(\lambda_f\) “small (on the order of 0.1 to 0.2)” — and why, if you must raise it, you should watch the format/correctness correlation for the decoupling described in the warning.

5. The math_verifier.py docstring promises to handle “Sets/tuples: {1, 2} == {2, 1}”, but neither try_numeric nor try_sympy actually parses set notation, so math_equivalent("{1, 2}", "{2, 1}") currently returns False. Implement a helper math_set_equivalent(pred, gold) in the chapter’s style that treats {...}, (...), or [...] as an unordered collection and compares members using the existing math_equivalent, then wire it into math_equivalent as an additional branch. Show that it makes {1, 2} == {2, 1} and {1/2, 0.25} == {0.5, 1/4} return True.

Solution

Reuse math_equivalent for element-wise comparison so each member gets full numeric/symbolic normalization. Parse the bracketed body, split on commas, then greedily match every predicted element to an unused gold element (unordered, with multiplicity).

import re

def try_collection(s: str):
    """
    Parse "{a, b, ...}", "(a, b, ...)" or "[a, b, ...]" into a list of
    raw element strings. Returns None if s is not bracketed.
    (Comma-split is flat: nested collections are out of scope.)
    """
    s = normalize_latex_number(s)
    m = re.match(r'^\s*[\{\(\[](.*)[\}\)\]]\s*$', s)
    if not m:
        return None
    body = m.group(1).strip()
    if not body:
        return []
    return [e.strip() for e in body.split(',')]


def math_set_equivalent(pred: str, gold: str) -> bool:
    """True iff pred and gold are the same unordered multiset of values."""
    ps = try_collection(pred)
    gs = try_collection(gold)
    if ps is None or gs is None:
        return False
    if len(ps) != len(gs):
        return False
    remaining = list(gs)
    for p in ps:
        for i, g in enumerate(remaining):
            if math_equivalent(p, g):   # recurse: full normalization per element
                del remaining[i]
                break
        else:
            return False                # no gold element matched this pred element
    return not remaining

Wire it into math_equivalent as a branch before returning False (after the symbolic strategy):

    # 4. Sets / tuples (unordered): "{1, 2}" == "{2, 1}"
    if math_set_equivalent(pred, gold):
        return True

    return False

Checks:

  • math_set_equivalent("{1, 2}", "{2, 1}"): ps=["1","2"], gs=["2","1"]; "1" matches "1", "2" matches "2" (order irrelevant), remaining empties -> True.
  • math_set_equivalent("{1/2, 0.25}", "{0.5, 1/4}"): "1/2" matches "0.5" via the numeric branch of math_equivalent, "0.25" matches "1/4", remaining empties -> True.

Because element comparison delegates to math_equivalent, mixed formats inside a set (fractions, decimals, LaTeX) are normalized for free; the only limitation is that flat comma-splitting does not handle nested collections such as {(1,2), (3,4)}.