Research deep dive

Two directions of learning

Can physical interaction produce a policy from which a useful predictive world model can later be recovered?

RESEARCH EXTENSION

Last updated · August 23, 2026

01 · Simple explanation

The question in plain language.

The same three components appear in opposite order. Both directions end at the same test: does the learned structure generalize?

02 · Visual walkthrough

Follow the mechanism.

FIG. 01Two directions of learning
Two directions of learning. Physical records flow toward a learned policy while robot behavior flows toward predictive structure, with both judged by one shared evaluation standard.

The reverse direction tests predictive usefulness; it does not assume that a learned policy contains explicit textbook physics.

03 · Full narrative

What the idea means and what it does not.

A conventional direction begins with a world model or simulator, uses it to train a policy, and then executes that policy in physical interaction.

The reverse direction begins with physical interaction, learns a policy, and asks whether that policy contains enough predictive structure to recover a useful model afterward.

Usefulness may mean supporting consequence prediction, synthetic trajectories, or transfer. Both directions must be tested on new objects, geometries, tasks, environments, or embodiments.

04 · Technical detail

A falsifiable path forward.

Experiment design

Compare a world-model-first pipeline with an interaction-first policy and recovered predictive model under matched tasks.

Baselines

Direct policy learning, conventional system identification, learned world models, and simulator-first training.

Metrics

Prediction accuracy, synthetic-trajectory usefulness, downstream control, transfer, and generalization.

Dependencies

Interaction histories with consequences, shared tasks, controlled simulator assumptions, and held-out environments.

Continue exploring

Place this idea in the larger program.