Reward as Observation

Learning Reward-Based Policies for Rapid Adaptation

A policy that sees only its recent rewards and actions never looks at the screen, so it keeps working when the screen changes. We train one in a cheap 2D simulator and deploy it, unchanged, in a 3D renderer and on a Stretch robot in scanned real homes.

Diagram: a policy pi receives reward r_t and outputs action a_t in a 2D top-down car racing game, and the same policy is reused on a 3D rendering of the same track.
Fig. 1The same policy drives in 2D and 3D. Its only inputs are reward rt and action at histories.

Overview

Method

Replace the observation with the reward

Source and target environments share state space, actions, dynamics, and reward. Only the observation differs, in any way at all: 2D to 3D, RGB to RGB-D, a new camera. We build a reward-based MDP by swapping the observation for a short history of reward and action pairs. That MDP is identical in source and target, so a policy trained in one runs in the other without modification.

Architecture diagram. Training phase: a chain of LSTM networks reads reward and action pairs, outputs value and action, trained with PPO plus a regression loss toward an observation-based expert's action. Testing phase: the reward comes either from the environment or from a network estimating reward from observations.
Fig. 3An LSTM reads reward/action pairs and is trained with PPO plus a guidance loss toward a state-based expert π*. At test time the reward comes from the environment (option 1) or from a learned reward estimator on target observations (option 2).
  1. Train in the source

    Train πrwd with PPO on reward/action history, regressing its actions toward an observation-based expert to keep exploration productive.

    cheap simulator · e.g. 2D Car Racing
  2. Deploy zero-shot

    Run the same checkpoint in the target. It never reads the target's observations, so nothing needs aligning.

    new observations · e.g. 3D render, Habitat-Sim
  3. Distill a student

    Use πrwd as a DAgger teacher to train an observation-based policy that needs no reward at inference.

    target pixels · CNN-LSTM student
Interactive · 2D Pointmass

Be the reward-based policy

The goal is hidden. You get the same input the policy gets: the reward after each move, r = exp(−3·d) where d is the distance to the goal. Probe, compare, and find it. Then watch a simple probing agent do the same.

0.000rt · step 0
a: none yet
Click here, then use the arrow keys or the pad.

The probing agent climbs one axis at a time using only reward changes, a hand-written stand-in for the learned LSTM policy.

Results

Train once, transfer across observations

Reward-based policies train across diverse environments

Across three seeds, reward-based policies reach 95%, 66%, and 70% of state-based performance in Pointmass, Cartpole, and Car Racing. Car Racing uses a dense reward, r = rdistance · rangle, built from distance to the next waypoint and heading error. The agent sometimes drifts off the centerline to gather more reward signal before correcting.

Four environment screenshots: Pointmass, Cartpole, top-down Car Racing, and a Habitat-Sim kitchen with a Stretch robot.
Fig. 2Pointmass, Cartpole (DeepMind Control Suite), 2D Car Racing (Gymnasium), and Habitat-Sim for transfer testing.
Three learning-curve plots of reward versus timestep. Green observation-based curves sit above orange reward-based curves in Pointmass, Cartpole, and Car Racing.
Fig. 4Learning curves, observation-based (upper bound) vs. reward-based.

Zero-shot transfer under palette swaps

We recolor every environment, including a pure noise background for Pointmass. The reward-based policy never sees the pixels, so its score barely moves.

Table II. Mean reward ± std over 50 trials.
ObservationPointmassCartpoleCar Racing
Original97.06 ± 1.36758.74 ± 62.79775.37 ± 225.27
Shifted95.96 ± 2.47755.45 ± 52.63715.93 ± 264.28
Grid of recolored observations: Pointmass on white, noise, green and gold backgrounds; Cartpole in four color schemes; Car Racing with magenta, teal and navy palettes.
Fig. 5Original observation (left column) and three deliberately difficult palette swaps.

From a 2D top-down view to a 3D camera

We built a 3D renderer with the same track layout and transition dynamics as Gymnasium Car Racing. The reward-based policy trained in 2D drives the 3D track directly, with behavior indistinguishable from the 2D expert.

Source2D Car Racing, top-down view.
Target3D Car Racing, driven by the DAgger student trained with the reward-based teacher.

A Stretch robot in scanned real homes

The same Car Racing checkpoint, with no retraining, drives a Stretch mobile robot along navmesh corridor tracks in two HM3D building scans in Habitat-Sim. Reward is computed from the robot's pose relative to recorded waypoints, which a real robot could get from localization. It succeeds in both scenes despite a different renderer and a different embodiment.

00801Bedroom to kitchen corridor.
00807Open living and dining area.
Two rows of seven frames showing a Stretch robot moving through a bedroom, hallway and kitchen, and through an open living and dining room.
Fig. 6Motion strips of the zero-shot deployment. Top: scene 00801. Bottom: scene 00807.

A teacher for pixel-based policies

Learning 3D Car Racing from pixels with a CNN-LSTM is slow, even on a log timestep axis. With the reward-based policy as a DAgger teacher, the pixel student reaches near-optimal driving in about 70,000 steps.

The upfront source training pays for itself twice. One reward-based policy serves every target that shares the reward and action spaces. And source simulators are usually light and parallel, while photorealistic targets cost far more per step.

Log-scale learning curve. DAgger in red reaches about 800 reward by 70K steps; pixel-based from scratch in green stays below 700 through 10M steps.
Fig. 73D Car Racing, log timestep scale: DAgger with the reward-based teacher vs. pixel-based learning from scratch.

Reward estimation succeeds where state estimation fails

When the target gives images but no reward, you can learn to estimate one. With 100,000 image samples in Cartpole, a learned five-dimensional state estimator drives the state-based policy down to random-policy level. A learned scalar reward estimator keeps the reward-based policy close to its source score.

Table III. Cartpole, mean ± std over 100 trials.
PolicyOriginal domainImage domain
State-based915.49 ± 2.24189.41 ± 12.54
Reward-based836.49 ± 4.46735.83 ± 65.48
Random–192.78 ± 36.91

Expert guidance matters most where failure ends the episode

Guidance helps a little in Pointmass and Cartpole, where poor exploration still earns reward. In Car Racing, leaving the track ends the episode and cuts off the reward signal. Guidance keeps episodes alive long enough to learn.

Three learning-curve plots comparing guidance in orange against no guidance in red. The gap is small in Pointmass and Cartpole and large in Car Racing.
Fig. 9Ablation: guidance vs. no guidance. The benefit grows with task difficulty.