LeMario: Training a JEPA World Model on Super Mario Bros

LeMario: Training a JEPA World Model on Super Mario Bros

I built LeMario from scratch to test if a JEPA model could learn Super Mario Bros dynamics without rewards. While the model predicted short-term futures accurately, it failed to navigate complex obstacles. My experiments revealed that predicting the game is not the same as learning how to make progress through it, offering a crucial lesson on the limits of current world models.

The model had learned to predict the game, but that did not mean it had learned how to make progress through it.
  1. enjeyw

    The author hints at this but it seems like one issues is that while JEPA is good at distinguishing between unpredictable noise and predictable features, the model has no way of assigning importance to different predictable features.

    So for a system where it’s very difficult to exactly reach the desired end state, the model needs to choose between (for example):

    - reaching a relatively achievable scene where 95% of the features in the latent are correct, which includes stuff like visible enemies, Mario’s position on the screen etc

    - reaching a far more difficult to access scene where there’s a bunch of differences in the actual level visuals, but theres a match on the latent for the tiny set of pixels in HUD that indicate you’ve hit the victory condition

    We obviously know that it’s not good enough to reach an early scene that looks similar to the victory condition but isn’t. The model doesn’t.

    In a sense, this is what the linear probe helps with - it allows us to re-weight the latent and say “actually, while the latent encodes many things about the world, the thing we really care about is the X position.”

    I’d be curious what happened if rather than planning actions on cross entropy of a final scene, the model just tried to find the actions that maximize the predicted X value of the probe.

  2. rsfern

    I think JEPA is super interesting, but I feel like this example highlights some of the challenges of long horizon planning. For one, chunking the planning stage into a bunch of intermediate goals seems really limiting, because a lot of what makes model based control interesting is that we don’t want to impose a solution strategy (because we want to solve problems we don’t know how to solve)

    Another thing that has been bothering me is that you have to write the goal in input space. That doesn’t align with all problems, for some problems there could be many different states that satisfy a goal. For Mario maybe it’s ok, but there’s some weirdness still, like should the goal state be Mario at the finish line of the level with a specific timer state in the frame header? What about optimizing the number of points?

    Also it’s interesting to think about how you would get Mario to reliably jump on koopas and goombas. IIUC JEPA models are usually trained with random rollouts, and then you’d handle this sort of intermediate goal in the planning optimizer? But that seems inefficient, and including some planning in the pretraining rollouts might be necessary to get enough relevant intermediate states. And then it starts feeling like reinforcement learning…

    I’d be happy to have a check on my intuition here, or pointers to interesting writing on these topics

    p.s. on topic, I liked the debugging strategies used in the blog post, that was my favorite part of the writeup

More from this day

2026-07-15