Tool-Using Language Agents
Reasoning–acting loops, learned API calls, reflective memory, executable skills and evaluation.
Read study →90-day research programme
A source-grounded research record spanning video generation systems, world-model formalisms, interactive-world architectures and embodied-agent control. Every claim exposes its evidence chain, epistemic level and unresolved uncertainty.
Latest record
Day 0 is not an empirical paper. It defines the evidence standard, publication pipeline and persistent research memory that govern the series.
Reasoning–acting loops, learned API calls, reflective memory, executable skills and evaluation.
Read study →The first theoretical layer: formal definitions, state assumptions, prior memory and the Phase 1 scope.
View programme →Primary and secondary evidence are separated; inference, speculation and contested propositions are never presented as fact.
Read methodology →Programme map
Days 1–15 · world-model formalism, POMDPs, latent dynamics and evaluation.
Days 16–35 · frontier systems, model cards and controllability analysis.
Days 36–55 · 3D geometry, persistence, drift and long-horizon consistency.
Days 56–75 · action conditioning, control, planning and sim-to-real.
Days 76–90 · benchmark design, testable hypotheses and roadmap.
01 · Formal substrate
The central distinction is between generating plausible observations and modelling a persistent, action-conditioned latent state.
World model: st+1 ∼ F(st, at), ot ∼ G(st)
A controllable generator becomes a candidate world model only when the latent state supports intervention, temporal persistence and counterfactual evaluation.
02 · Epistemic infrastructure
Claims grounded in an opened paper, official technical report, repository or first-party project documentation.
Claims supported by reliable secondary analysis while preserving the distinction from primary evidence.
Original synthesis and speculation remain explicitly labelled, testable and separate from reported results.
03 · Measurement
Visual quality alone is insufficient. The programme evaluates whether generated futures remain coherent under intervention.
| Axis | Operational question | Failure signal |
|---|---|---|
| Temporal coherence | Does state identity persist across long rollouts? | Object drift, identity resets, irreversible scene changes. |
| Causal control | Does an action produce the expected intervention? | Correlated motion without counterfactual validity. |
| Geometry | Is unseen structure conserved when viewpoint changes? | View-dependent hallucination and topology collapse. |
| Physics | Do contact, inertia and occlusion remain consistent? | Violation of constraints outside the training distribution. |
04 · Persistence
What information survives the encoder bottleneck, and is it sufficient for future prediction and action selection?
Objects, geometry and affordances require representations that survive occlusion, camera motion and long horizons.
Every rollout accumulates uncertainty; evaluation must expose when prediction error becomes a planning failure.
05 · Agency
The decisive test is not whether a model can render a plausible frame, but whether an agent can use its predictions to select actions.
Control loop: observe → infer belief state → imagine candidate trajectories → score outcomes → act → update memory.
The analysis tracks action spaces, policy interfaces, model-predictive control, imagination-based reinforcement learning and sim-to-real transfer.
06 · Research agenda
Which latent or 3D representations preserve the variables required for intervention rather than appearance matching?
How can causal world-model competence be measured independently of pixel-level video quality?
What memory, control and safety guarantees are required before simulated futures can guide real agents?