The new pretraining method replaces target encoders and reconstruction machinery with a single encoder trained using invariance and SIGReg. This end-to-end approach achieves performance equivalent to V-JEPA 2 while reducing compute costs by 5.6 to 20.8 times.

LeVJEPA utilizes a stable recipe to maintain appearance and motion understanding, even when randomly dropping 95% of video tokens. The method is open source and removes the need for masked predictions, stop-gradients, or teacher-student schedules.

Sign in to suggest edits

Key sources

  1. SOURCE@lukaskuhn77“No target encoder, no masked prediction, no stop-gradient or teacher-student schedule”x.com
  2. SUPPORT@randall_balestr“open source + reproducible”x.com
  3. SUPPORT@askalphaxiv“randomly dropping 95% of video tokens, preventing representation collapse”x.com
  4. SUPPORT@ylecun“A new Pareto frontier in video pretraining”x.com
Markdown