Compare a world-model-first pipeline with an interaction-first policy and recovered predictive model under matched tasks.
Research deep dive
Two directions of learning
Can physical interaction produce a policy from which a useful predictive world model can later be recovered?
RESEARCH EXTENSIONLast updated · August 23, 2026
01 · Simple explanation
The question in plain language.
The same three components appear in opposite order. Both directions end at the same test: does the learned structure generalize?
02 · Visual walkthrough
Follow the mechanism.

The reverse direction tests predictive usefulness; it does not assume that a learned policy contains explicit textbook physics.
Reserved for an approved motion explanation. The walkthrough and caption above provide the complete text equivalent.
03 · Full narrative
What the idea means and what it does not.
A conventional direction begins with a world model or simulator, uses it to train a policy, and then executes that policy in physical interaction.
The reverse direction begins with physical interaction, learns a policy, and asks whether that policy contains enough predictive structure to recover a useful model afterward.
Usefulness may mean supporting consequence prediction, synthetic trajectories, or transfer. Both directions must be tested on new objects, geometries, tasks, environments, or embodiments.
04 · Technical detail
A falsifiable path forward.
Direct policy learning, conventional system identification, learned world models, and simulator-first training.
Prediction accuracy, synthetic-trajectory usefulness, downstream control, transfer, and generalization.
Interaction histories with consequences, shared tasks, controlled simulator assumptions, and held-out environments.