BoNVoyageLearning Better Rewards
without Ranking

A reward model should remain useful when its own scores drive optimization. BoNVoyage learns rewards by fitting the policy they induce to preferred demonstrations.

01 / THE PROBLEM

The problem: reward over-optimization

Good at ranking preference data ≠ a good reward signal.

As RL optimizes a learned reward, the policy concentrates on responses the reward model scores highly, including responses that exploit its errors. Reward can keep increasing while actual response quality declines.

Bradley–Terry learns from fixed response pairs. It does not directly train the reward for the response distribution that optimizing it creates.

How do we train reward models for their downstream use in RL?

02 / THE METHOD

Train the induced policy

Maximize the likelihood of preferred responses under the policy induced by the reward model.

πϕ(y∣x)Policy induced bythe reward model↓  ∝  pLM(y∣x)↑Base LMexp⁡ ⁣(  reward⁡ϕ(y,x)Reward model↓ / ⁣β↑Regularizerstrength)\overset{\color{#ff9b51}\substack{\text{Policy induced by}\\\text{the reward model}\\\downarrow}}{\pi_\phi(y\mid x)}\;\propto\;\underset{\color{#ff9b51}\substack{\uparrow\\\text{Base LM}}}{p_{\mathrm{LM}}(y\mid x)}\exp\!\bigl(\;\overset{\color{#ff9b51}\substack{\text{Reward model}\\\downarrow}}{\operatorname{reward}_\phi(y,x)}\,/\!\underset{\mathclap{\color{#ff9b51}\substack{\uparrow\\\text{Regularizer}\\\text{strength}}}}{\beta}\bigr)

The base language model supplies the reference distribution; β controls regularization. The reward determines which responses become more likely.

03 / How it works

  1. 01

    Start with a preferred response.

    Initialize a short Markov chain at a good response selected from the base LM, y⁺.

  2. 02

    Use our MCMC method with the base LM.

    Rescore a reused candidate pool with the current RM. Accept proposals using reward differences.

    α(y′,y)=min⁡ ⁣{1,exp⁡ ⁣(rϕ(y′,x)−rϕ(y,x)β)}\alpha(y',y)=\min\!\left\{1,\exp\!\left(\frac{r_\phi(y',x)-r_\phi(y,x)}{\beta}\right)\right\}
  3. 03

    Get the chain’s final sample.

    Use it as the negative example y⁻ for the reward update.

    Likelihood gradient∇ϕlog⁡πϕ(y+∣x)≈1β[∇ϕreward⁡ϕ(y+)−∇ϕreward⁡ϕ(y−)]\nabla_\phi\log\pi_\phi(y^+\mid x)\approx\frac{1}{\beta}\bigl[{\color{#ffad66}\nabla_\phi\operatorname{reward}_\phi(y^+)}-{\color{#8acfff}\nabla_\phi\operatorname{reward}_\phi(y^-)}\bigr]

04 / A REWARD THAT STAYS USEFUL

A reward that stays useful

We study the method with 4B reward models and Qwen3-1.7B policies in mathematics and science. We evaluate how the reward behaves during downstream RL and BoN selection.

Paper Figure 1 compares math accuracy and reward-correctness correlation during RL; BoNVoyage remains stronger than the Bradley–Terry baselines later in training.
Off-policy BTOn-policy BTBoNVoyageRLVR
Figure 1 · RL on math data (DeepScaleR). Qwen3-1.7B-Base optimized with 4B reward models. Left: training accuracy; dashed RLVR uses ground-truth rewards as a reference upper bound. Right: Pearson correlation between reward scores and response correctness. RLVR is omitted here because its correlation is 1 by construction.

05 / DOWNSTREAM RL

Downstream RL
in two settings

MATH-500 1.7B Base · 4B RM

Model / rewardAccuracy (%)
Base model25.2 ± 1.0
On-policy BT63.4 ± 1.8
Skywork (Off-policy BT)63.0 ± 1.8
BoNVoyage67.7 ± 1.8

Science 1.7B Science SFT · 4B RM

MethodSciQSciKnowEvalARC-EasyARC-CMMLU-Sci
Science SFT82.5 ± 0.967.3 ± 1.279.5 ± 0.665.7 ± 1.046.6 ± 1.5
On-policy BT91.1 ± 0.772.2 ± 1.291.5 ± 0.478.7 ± 0.957.4 ± 1.6
Skywork
(Off-policy BT)
90.4 ± 0.873.1 ± 1.290.0 ± 0.576.7 ± 1.057.4 ± 1.7
BoNVoyage92.3 ± 0.774.1 ± 1.293.1 ± 0.481.3 ± 0.959.3 ± 1.7

Accuracy (%) ± reported uncertainty. Skywork: Skywork-Reward-V2-Qwen3-4B. BT is early-stopped on validation before over-optimization.

06 / BoN

Stronger RM with BoN

Paper Figure 2 shows BoN accuracy on MATH-500 and SciKnowEval across sampling budgets up to 512.
Select the highest-reward response from n base-LM samples. Plots reproduced directly from the paper.

ABSTRACT

Abstract

Reward models (RMs) play a central role in reinforcement learning from human feedback (RLHF), yet they are typically trained with a Bradley–Terry (BT) objective that ranks response pairs under a fixed text distribution. This objective does not match how RMs are used in RLHF: as training progresses, the policy increasingly concentrates on responses that receive high reward under the RM itself. But what matters is not whether the RM ranks fixed pairs correctly, but whether it remains reliable on the responses that its own signal makes more likely under the policy. We develop BONVOYAGE, a method for training RMs directly for this downstream role. Instead of training a pairwise ranker, BONVOYAGE maximizes the likelihood that the RM-induced RLHF-optimal policy assigns to preferred responses. Estimating the gradient of this objective requires samples from the optimal policy induced by the learning RM. We obtain these samples through test-time alignment: given the base LM and the RM, we use Markov chain Monte Carlo to sample from the induced policy. To make this practical, we use the base LM as an independent proposal distribution, reuse candidate responses across optimization steps, and use contrastive divergence with short chains initialized at preferred responses. We evaluate BONVOYAGE by training 4B RMs for QWEN3-1.7B-BASE in mathematics and our supervised finetuned model QWEN3-1.7B-SCIENCE in science. Compared with BT baselines, BONVOYAGE yields stronger downstream policies under RLHF and stronger rerankers under best-of-n sampling across MATH-500, GSM8K, SciQ, SciKnowEval, ARC-Challenge, and MMLU-Science. On verifiable math tasks, BONVOYAGE also shows substantially more robustness to reward over-optimization, preserving alignment between RM scores and ground-truth correctness later into RLHF training.

CITATION

Build on this work.

Read the paper
BibTeX
@misc{faria2026bonvoyage,
  title = {BoNVoyage: Learning Better Rewards without Ranking},
  author = {Faria, Gonçalo and Yang, Guang and Liu, Alisa and Smith, Noah A.},
  year = {2026},
  url = {https://openreview.net/forum?id=W5Z8iFhtik}
}