Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures

José Lucas De Melo Costa, Seong Woo Ahn, Fabrice Popineau, Arpad Rimmel, Bich-Liên Doan

Université Paris-Saclay, CNRS, CentraleSupélec, LISN

NeurIPS 2026, Main Track (poster), Paris, 8–12 December 2026.

Whether a JEPA collapses is decided early, direction by direction, by one ratio: drive over decay.

The mechanism in 33 seconds. The cloud follows the exact solution of the paper's linear model.

In sixty seconds

A JEPA learns by predicting the representation of a hidden part of the input from the visible part. Sending every input to the same point also satisfies that objective, and whether training ends there is decided in its first steps.

For linear networks we prove that each direction of the representation is pulled out by a drive, which comes from the agreement between the two views, and pulled back by a decay, which comes from the predictor and the data. The direction is learned when μ = drive / decay exceeds 1.

Counting the directions with μ > 1 gives a rank budget for the representation. Across more than 800 Tabular-JEPA configurations the transition between learning and collapse lines up with the measured μ = 1 contour, and in the linear model seven known anti-collapse mechanisms reduce to doses on μ.

The ratio suggests starting the predictor's attention at the identity. With this ResidualPred, I-JEPA gains 4 to 5.6 points of linear-probe accuracy on CIFAR-10, CIFAR-100, STL-10 and ImageNet-100, and 6.5 points on an ImageNet-1k pilot at a matched budget.

Phase diagram over predictor scale and context ratio: green where the drive wins (learning), red where the decay wins (collapse), separated by the contour mu = 1.
Figure 1 of the paper. Learning where the drive wins (μ = γ/σ > 1), collapse where the decay wins (μ < 1).
Interactive tutorial

Hold the knob

A twenty-minute tutorial that you run in your browser. Make a JEPA collapse, find the ratio that decided it, map the phase boundary, then fix it with the identity. Every widget solves the paper's linear model exactly, and four Python cells let you change the experiment.

Open the tutorial →
Jose Costa

Open to collaborations

I work on the training dynamics of joint-embedding predictive architectures, on conformal and distribution-free uncertainty, and on tabular foundation models.

PhD (CIFRE) at LISN, CentraleSupélec, Université Paris-Saclay, until autumn 2027. Happy to talk about joint work, research visits and what comes after the PhD.

Other recent work

  • Single-Pass Conformal Cross-Modal Anomaly Screening with Tabular Foundation Models

    Oral presentation at MICCAI 2026 @ MultiTab Workshop, Strasbourg.

  • Mitigating Convergence Collapse in Fixed-Target Anomaly Detectors via Kernel-Anchored Locality Regularization

    Oral presentation at CIKM 2026, Rome.

  • Knowledge-Informed Local Causal Discovery of Optimal Adjustment Sets

    CIKM 2026, Rome.

  • T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular Data

    ICLR 2025. *Joint first authors.