CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization

A latent world model that splits its representation into controllable and uncontrollable subspaces — staying robust to visual distractors where prior JEPA world models collapse.

1Georgia Institute of Technology 2Georgia Tech Research Institute

Controllability factorization

CF-JEPA builds on the LeWM world model. The encoder \(f\) maps an observation \(x_t\) to a latent \(z_t=[z^c_t, z^u_t]\). An action-conditioned predictor \(g\) advances the controllable part, while an action-free predictor \(h\) advances the uncontrollable part. An inverse-dynamics head \(\psi\) forces \(z^c\) to carry action-relevant information; an adversarial head \(\phi\) (with a gradient-reversal layer) prevents \(z^u\) from predicting actions. A SIGReg term keeps the latent from collapsing.

CF-JEPA method overview: encoder, controllable and uncontrollable predictors, inverse and adversarial dynamics heads, and the four training losses.
Controllable subspace \(z^c\in\mathbb{R}^{128}\) — used for control Uncontrollable subspace \(z^u\in\mathbb{R}^{64}\) — absorbs distractors

The full objective combines forward prediction, inverse dynamics on \(z^c\), an adversarial term on \(z^u\), and the SIGReg regularizer:

$$\mathcal{L}=\mathcal{L}_{\text{fwd}} + \alpha\,\mathcal{L}_{\text{inv}} + \beta\,\mathcal{L}_{\text{adv}} + \gamma\,\mathcal{L}_{\text{SIGReg}}$$

with \(\alpha=1\) and \(\gamma=0.09\) fixed for all tasks, and \(\beta\) the single per-task tuned weight. At planning time, model-predictive control (CEM) optimizes actions using only the controllable latent \(z^c\).

Tasks & visual distractors

We evaluate on four control tasks — DM Control Reacher, PushT, OGBench Cube, and TwoRoom — plus a simulated ManiSkill robot reach. Robustness is stress-tested by adding \(N\) moving colored rectangles that drift along the image edges without occluding the task.

The four primary evaluation tasks: Reacher, PushT, OGBench Cube, TwoRoom.
Figure 1. Primary tasks: DM Control Reacher, PushT, OGBench Cube, and TwoRoom.
The same tasks with example moving-rectangle distractors added.
Figure 3. The same tasks with example distractors (\(N=2\)).

Robust where baselines collapse

Under nominal conditions CF-JEPA is competitive with LeWM and SMWM. Once distractors are added, LeWM collapses on all tasks and SMWM collapses on Reacher — CF-JEPA is the only model that never collapses. Values are success rate (higher is better); CF-JEPA is reported as mean ± std over three seeds.

Table II. Nominal task performance (no distractors).
TaskLeWMSMWMCF-JEPA
Reacher0.860.660.61 ±0.03
PushT0.960.830.78 ±0.04
OGBench Cube0.740.840.81 ±0.03
TwoRoom0.870.990.95 ±0.04
Table III. Performance under the standard distractor setting (\(N=2\)).
TaskLeWMSMWMCF-JEPA
Reacher0.07 ±0.050.10 ±0.030.48 ±0.12
PushT0.01 ±0.010.86 ±0.020.87 ±0.03
OGBench Cube0.42 ±0.020.84 ±0.010.85 ±0.02
TwoRoom0.33 ±0.030.98 ±0.050.95 ±0.03
Table IV. Task performance as the number of distractors \(N\) increases. Best per row in orange.
Task\(N\)LeWMSMWMCF-JEPA
Reacher20.080.10 ±0.030.48 ±0.12
PushT20.010.86 ±0.020.87 ±0.03
16—0.36 ±0.360.76
32—0.57 ±0.010.68 ±0.06
64—0.27 ±0.250.57 ±0.01
OGBench Cube20.410.84 ±0.010.85 ±0.02
32—0.80 ±0.020.81 ±0.01
64—0.70 ±0.050.73 ±0.01
128—0.73 ±0.050.70 ±0.02
TwoRoom20.34 ±0.050.98 ±0.050.95 ±0.03
8—1.00 ±0.000.97 ±0.01
16—0.60 ±0.280.94 ±0.00
32—0.46 ±0.120.82 ±0.00

Planning rollouts under distractors

Rollouts on Reacher with distractors (\(N=2\)). CF-JEPA reaches the target; both LeWM and SMWM collapse and lose track of the arm.

CF-JEPA Reacher rollout with distractors.
CF-JEPAreaches target
SMWM Reacher rollout with distractors.
SMWMcollapses
LeWM Reacher rollout with distractors.
LeWMcollapses

Increasing distractor density on PushT (\(N=16\)): CF-JEPA holds up, while SMWM fails.

CF-JEPA PushT rollout at N=16 distractors.
CF-JEPA · PushT at \(N=16\)solves
SMWM PushT rollout failing at N=16 distractors.
SMWM · PushT at \(N=16\)fails

What each subspace remembers

Decoding frozen encoders (trained with distractors) back to pixels on clean Reacher frames shows the factorization at work: CF-JEPA's controllable latent \(z^c\) reconstructs the arm, its uncontrollable latent \(z^u\) does not, and both baselines have lost the agent entirely to collapse.

CF-JEPA controllable latent reconstruction of the Reacher arm.
CF-JEPA \(z^c\) — arm reconstructedretains agent
CF-JEPA uncontrollable latent, unable to reconstruct the arm.
CF-JEPA \(z^u\) — arm blurred away
SMWM reconstruction, collapsed.
SMWMcollapsed
LeWM reconstruction, collapsed.
LeWMcollapsed
Figure 5. Only the controllable latent of CF-JEPA retains the agent under distractor training.

We quantify collapse with the participation ratio of the latent, computed from the eigenvalues \(\lambda_i\) of the embedding covariance over 256 random frames — a higher value means variance is spread across more dimensions, while total collapse drives it to 1:

$$\mathrm{PR}=\frac{\left(\sum_{i=1}^{d}\lambda_i\right)^{2}}{\sum_{i=1}^{d}\lambda_i^{2}}$$

Table VI. Latent participation ratio on Reacher (higher = more spread; collapse → 1).
ConditionLeWMSMWMCF-JEPA
Nominal48.2 ±0.83.71 ±0.0742.3 ±0.1
Distractor (\(N=2\))2.24 ±0.231.00 ±0.0112.9 ±0.35

Perception-dependent ManiSkill reach

Simplified ManiSkill reach task.
Figure 4. Simplified ManiSkill reach used for distractor evaluation.
Table V. Success rate at increasingly tight goal thresholds (\(N=16\) distractors).
Threshold (m)LeWMSMWMCF-JEPA
0.06 (easiest)0.55 ±0.040.79 ±0.050.83 ±0.06
0.050.51 ±0.050.62 ±0.040.72 ±0.08
0.040.44 ±0.030.58 ±0.020.64 ±0.10
0.030.36 ±0.030.38 ±0.020.53 ±0.14
0.0250.32 ±0.060.30 ±0.040.48 ±0.17
0.02 (hardest)0.21 ±0.050.22 ±0.040.35 ±0.12