← survey

The debates

Latent prediction vs. pixel generation

The field’s central fault line. LeCun’s position: predicting raw pixels wastes capacity on unpredictable, irrelevant detail; a world model should predict abstract representations in latent space — which, per the survey literature (Ding et al.), makes learning “more efficient and robust.” The JEPA family is this position implemented (see latent prediction).

The generative counter-position: scale video generation far enough and a world simulator falls out — this was Sora’s framing. The verified critique (see video-generative): original Sora neither conditions on actions nor reproduces physics reliably, and DeepMind’s Physics-IQ found its visual plausibility dissociated from physical understanding.

The debate is not settled — it is converging. Cosmos 3 is generative and action-conditioned in one omnimodal system; V-JEPA 2-AC gives the latent camp real robot control. Both camps are drifting toward the “unified multimodal world model” the surveys forecast. Whether Meta continues to position JEPA as a distinct non-generative path is an open question we track.

Do LLMs have world models?

The second debate: whether next-token prediction over text induces an internal model of the world, and whether that (rather than explicit world modeling) is a sufficient path to general intelligence. Evidence runs both directions — the Othello-GPT probing line suggests sequence models induce board-state representations; skeptical reviews find the evidence for abstract, robust world models in LLMs weak. (Claims here are contested; we hold them to the same verification bar as everything else, and detailed verification is in progress.)

The debate now has a video-native chapter: DeepMind’s “Video models are zero-shot learners and reasoners” (arXiv:2509.20328) argues Veo 3 exhibits emergent reasoning — 78% zero-shot maze solving vs. Veo 2’s 14% — via “chain-of-frames”, the visual analog of chain-of-thought. If frame-by-frame generation is a reasoning substrate, the world-models-vs-LLMs question becomes which modality’s prediction objective builds the better internal model — and both camps now have billion-dollar bets riding on their answer (see companies & labs: AMI for latents, everyone else for pixels).

Evaluation as a debate

The third argument is about measurement itself: the field has no agreed way to score “having a good world model.” Physics consistency, long-horizon coherence, action fidelity, and downstream control success are all proxies, and the surveys name fragmented evaluation as a principal open problem — see evaluation and open problems.


← Driving world modelsEvaluating world models →