Edit .env with local dataset and output roots.

September 14, 2026 · View on GitHub

INTACT: intent to action

INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

Train a world model to answer the control query it will receive at deployment.

End-to-end JEPA world modeling for goal-conditioned robot control without broad action search.

Junhan Sun1,4  ·  Hao Zhao2,4,†  ·  Guofeng Zhang1,3,†

1State Key Laboratory of CAD&CG, Zhejiang University
2Institute for AI Industry Research (AIR), Tsinghua University
3InSpatio  ·  4RoboParty Lab
Corresponding authors

Paper on arXiv Live project page Code available Models on Hugging Face

Join the bilingual INTACT World Model Community

Why INTACT?  ·  Installation  ·  Training  ·  Evaluation  ·  Method Notes  ·  Results  ·  中文

News

  • [2026-09-14] Evaluation and model update: Fixed duplicate previous-action insertion during actor-backed evaluation. Under the corrected causal-history protocol, task-specific E1 INTACT reaches 95.61 +/- 0.59% Official Direct macro SR and 96.58 +/- 0.44% with optional Guarded A. We also released history-free INTACT checkpoints and results as a separate ablation.
  • [2026-08-06] Code release: Open-sourced training and evaluation code, task-specific and shared-encoder configurations, reproducibility tools, and model documentation.
  • [2026-07-28] Project release: Released the paper, project page, and project film.

Project Film

Watch the INTACT project film

Download the share edition with project QR  ·  Open the interactive project page

INTACT v31 teaser: matched control, action-aligned representation, and search-free inference

A strong representation keeps the information that matters intact.
INTACT does exactly that, turning LeWM into a stronger world model.

Why INTACT?

We introduce INTACT (INtent-To-ACTion), an end-to-end JEPA that turns action-labeled, reward-free trajectories into a deployable intent-to-action interface. The name captures both the structure we impose and the information we preserve:

  • Isomorphic between predictor graphs. Local and goal motion-intent calls use the same four-slot input grammar and the same parameters.
  • Isomorphic between supported families. Local and goal intent families correspond through the action-law semantics induced by that shared predictor, not through pointwise latent equality.
  • Intact from RGB evidence to latent intent. End-to-end action gradients retain action-effective visual information while suppressing nuisance that is unrelated to motion intent.
  • Intact from intent families to action-law families. The shared predictor preserves the supported family correspondence all the way to direct action readout.

TL;DR

Forward world models answer "what will happen if I execute this action?" Yet deployment asks the inverse question: "which action realizes this intent?" CEM and MPPI answer it by numerically searching over many candidates, leaving training and inference without a learned semantic correspondence. The result is often a predictor plus an action searcher, rather than a self-consistent intent-action model.

INTACT learns that missing correspondence end to end. One conditional operator maps both observed physical change and deployable goal intent to an action law. Its conditional mean is a zero-search controller; sampling is retained only for diversity or optional local verification. No frozen encoder, extra policy-training stage, or globally linear latent dynamics is required.

1 epoch training  ·  95.61% Direct macro  ·  0 search  ·  3.9-4.8 ms latency

91.22% Shared E5 Direct  ·  96.58% Guarded  ·  23.44x fewer candidates

Current Results and Models

SettingInferenceMacro SRModel / details
Task-specific E1 INTACTDirect, zero search95.61 +/- 0.59INTACT
Task-specific E1 INTACTGuarded A 128x396.58 +/- 0.44Audited results
Shared-encoder E5 INTACTDirect, zero search91.22 +/- 0.51INTACT-unified
History-free E1 ablationDirect, zero search94.25 +/- 0.08INTACT-no-previous-action

The history-free row is a separate ablation and should not be confused with the optional Guarded A result. Exact task values, variances, and evaluation contracts are recorded in Audited Results.

Implementation and Reproduction

The released implementation includes task-specific training, four-task shared-encoder training, Direct/CEM/Guarded-A inference, Official LeWM and CLEAR-LeWM v0.8 scoring adapters, checkpoint manifests, and a checkpoint-compatible paper evaluation runtime. Smoke mode exercises the real data, model, optimizer, and checkpoint path; it is a code-path check, not an accuracy claim.

Installation

The locked CUDA environment was validated on Ubuntu 22.04, Python 3.10, PyTorch 2.6.0, and CUDA 12.4:

git clone https://github.com/zju3dv/INTACT-JEPA.git
cd INTACT-JEPA
bash scripts/install.sh cu124
source .venv/bin/activate
cp .env.example .env
# Edit .env with local dataset and output roots.
source scripts/fleet_env.sh
"$INTACT_PYTHON" scripts/verify_install.py --require-cuda

Use bash scripts/install.sh cpu only for import and configuration checks on a machine without an NVIDIA GPU. Full training and evaluation require CUDA. See Installation for system packages, manual setup, data conversion, and first-run diagnostics.

Data

INTACT uses the official LeWM datasets. Existing data are reused in place and are never downloaded implicitly by the training scripts. Configure this layout through .env:

$LOCAL_DATASET_DIR/
└── datasets/
    ├── pusht_expert_train.lance
    ├── ogbench/cube_single_expert.h5
    ├── reacher.h5
    └── tworoom.h5

The public PushT archive is HDF5. Convert it once to the Lance training layout after placing pusht_expert_train.h5 in datasets/:

"$INTACT_PYTHON" -m stable_worldmodel.cli convert \
  pusht_expert_train pusht_expert_train.lance \
  --source-format hdf5 --dest-format lance
"$INTACT_PYTHON" scripts/verify_data.py
bash scripts/check_fleet.sh

Smoke tests

# One task, one GPU, one optimizer step.
CUDA_VISIBLE_DEVICES=0 bash scripts/train.sh goal pusht \
  --smoke --run-name smoke_goal_pusht

# Shared encoder, one task per GPU, one synchronized optimizer step.
CUDA_VISIBLE_DEVICES=0,1,2,3 bash scripts/train_multitask.sh \
  --smoke --run-name smoke_multitask_goal

Run directories are immutable by default: an existing log or non-empty output directory causes an actionable failure instead of silent overwriting.

Training

Training always uses Math SDPA.

Single-Task Training

The task-specific paper setting trains one model per task for one full-data epoch:

CUDA_VISIBLE_DEVICES=0 bash scripts/train.sh goal pusht \
  --run-name intact_goal_pusht_s3072_e1 \
  seed=3072 trainer.max_epochs=1

Replace pusht with cube, reacher, or tworoom, and replace goal with waypoint for the matched coordinate control.

Each effective eight-frame window trains all seven physical and all seven goal transitions. On every index, the same demonstrated action is evaluated under the attached successor condition and the detached deployment-goal condition. An interior clip uses its true preceding action block; only unavailable history before a real episode boundary is raw-zero padded and then normalized.

Joint Multi-Task Training

The four-task E5 setting uses the same seven-local/seven-goal objective in four processes, one task per GPU, with one shared encoder/projector and task-specific Forward/action heads:

CUDA_VISIBLE_DEVICES=0,1,2,3 bash scripts/train_multitask.sh \
  --run-name intact_multitask_goal_s3072_e5

Its released default is fused AdamW, constant lr=5e-4, weight decay 1e-3, SIGReg 0.03, five epochs, Math SDPA, and batch 256 per task. Do not silently reduce the batch when reporting a paper-protocol reproduction.

Both launchers print all artifact paths and emit machine-readable progress at step 1, every 100 steps, and the final step:

TRAIN_PROGRESS={"epoch": 1, "global_step": 100, "loss": ..., "lr": ..., "eta_seconds": ...}
MULTITASK_PROGRESS={"task": "pusht", "rank": 0, "global_step": 100, ...}

Evaluation

INTACT has one policy-evaluation entrypoint. MODE selects the controller; it does not select a different history protocol:

ModeActorSearchRole
directyesnoneNative search-free controller
cemnobroad CEMActor-disabled representation/control baseline
guarded_ayeslocal CEM 128x3H5/RH5, sigma=0.25, top-k 16 around Direct
bash scripts/eval.sh direct pusht \
  intact_goal_pusht_s3072_e1/weights_epoch_1.pt 42 100
bash scripts/eval.sh cem pusht \
  intact_goal_pusht_s3072_e1/weights_epoch_1.pt 42 100
bash scripts/eval.sh guarded_a pusht \
  intact_goal_pusht_s3072_e1/weights_epoch_1.pt 42 100

All actor-backed modes use the same causal continuation contract. At a sampled start step t, the initial actor input is rows[t-5:t]; missing rows before the true episode boundary are raw-zero padded. After reset, the history shifts in actions actually executed by the controller. The current dataset action row[t] and target action are never exposed.

The command above uses Official LeWM scoring. CLEAR-LeWM v0.8 is a different benchmark protocol, not a different INTACT history mode. Install its pinned environment and invoke the scoring adapter directly:

export CLEAR_LEWM_ROOT=/path/to/CLEAR-LeWM-v0.8
export CLEAR_LEWM_PYTHON="$CLEAR_LEWM_ROOT/.venv/bin/python"
MANIFEST="$CLEAR_LEWM_ROOT/manifests/v0.8/pusht/moderate-seed42-n100.json"
CHECKPOINT=intact_goal_pusht_s3072_e1/weights_epoch_1.pt

"$CLEAR_LEWM_PYTHON" clear_eval.py \
  --clear-root "$CLEAR_LEWM_ROOT" \
  --manifest "$MANIFEST" \
  --policy "$CHECKPOINT" \
  --output results/pusht-clear-direct.json \
  --mode direct

Use --mode pure_cem or --mode guarded_a for the corresponding CLEAR run. Official and CLEAR results must be reported separately because their reset manifests and success criteria differ.

Paper checkpoints

The checked-in manifests define six controlled shared-encoder E5 cells, three training seeds, and four task shards per seed (72 checkpoints). The immutable paper-e5-goal-v1 model revision is publicly hosted at INTACT-JEPA/INTACT. These commands anonymously download and verify one cell or the full matrix:

bash scripts/download_paper_checkpoints.sh waypoint_intact all
bash scripts/download_paper_checkpoints.sh matrix all
bash scripts/eval_paper_matrix.sh goal_intact pusht 3072 42 100

Paper checkpoints use the bundled compatibility runtime in paper_runtime/. Exact cell mappings, expected paper scores, hash verification, and compatibility boundaries are documented in Paper Checkpoints.

Reproduction record

For every reported run, retain the source commit, resolved configuration, dataset and protocol, task, training/evaluation seeds, epoch, batch per task, inference mode, SDPA backend, checkpoint hash, completion metadata, and machine-readable evaluator output. The launchers record these fields where available. The complete reporting requirements remain normative in the Reproducibility Contract.

Core Insight: One Input Grammar, Two Intent Instances

INTACT has one predictor input form. For either intent instance mtm_t, it uses

xt(mt)=[zt,mt,ztmt,A(at1)],Gη(xt(mt))=pη(atxt(mt)).x_t(m_t)=\big[z_t,m_t,z_t\odot m_t,A(a_{t-1})\big], \qquad G_\eta\left(x_t(m_t)\right) =p_\eta(a_t\mid x_t(m_t)).

Only the value and gradient role of mtm_t change:

mtlocal=zt+1zt,mtgoal=sg(zg)zt.m_t^{\mathrm{local}}=z_{t+1}-z_t, \qquad m_t^{\mathrm{goal}}=\mathrm{sg}(z_g)-z_t.

The local instance uses a realized successor to ground which physical change produced the demonstrated action. The goal instance presents the same operator with the intent available before acting. Both come from the same demonstration and are supervised against the same correct ata_t, but each supervised conditional remains one triplet (zt,mt,at)(z_t,m_t,a_t) with one endpoint and one proper NLL:

LI2A=λlocal[logpη(atxt(mtlocal))]+λgoal[logpη(atxt(mtgoal))].\mathcal L_{\mathrm{I2A}} =\lambda_{\mathrm{local}}[-\log p_\eta(a_t\mid x_t(m_t^{\mathrm{local}}))] +\lambda_{\mathrm{goal}}[-\log p_\eta(a_t\mid x_t(m_t^{\mathrm{goal}}))].

There is no direct loss between the two endpoints or their latent displacements. Instead, the shared likelihood creates a conditional action quotient: at a fixed state, intents are equivalent when they induce the same expert action law. Under a task-appropriate tolerance, nearby predictions a^t(1)\hat a_t^{(1)}, a^t(2)\hat a_t^{(2)}, and the demonstrated ata_t can therefore belong to the same action-equivalence neighborhood. This distributional view makes direct control less sensitive to small prediction errors and helps limit closed-loop drift, while the forward JEPA keeps the richer world information needed for prediction.

Method

The physical successor remains attached to ground reachability; the future goal is a stop-gradient deployment anchor. INTACT aligns the two supported condition families through the actions they induce, without matching their endpoints, imposing globally linear latent dynamics, freezing the encoder, or adding a phase-2 controller. The goal likelihood alone is a goal-conditioned imitation objective; full INTACT is its shared, end-to-end coupling with the attached physical likelihood and forward JEPA.

INTACT retains the forward JEPA and adds one action-law predictor with a matched input grammar for physical and deployable intents. The key construction is not two unrelated auxiliary losses: both calls share parameters and a proper action likelihood, while their upstream gradient routes remain deliberately asymmetric.

ComponentRole
Forward predictorPreserves latent dynamics, contacts, topology, and visual information needed for rollout.
Physical intent callUses the observed successor with attached gradients to preserve action-recoverable change.
Goal intent callUses a detached future goal to train the same interface on a condition available before acting.
Matched interactionPairs first-order intent mtm_t with the state-intent feature ztmtz_t\odot m_t.
Direct controllerEmits an action chunk with no candidate search or terminal latent-cost call.
Guarded local verificationOptionally refines the coherent Direct plan with a small local CEM budget.

The four-domain model shares one visual encoder and keeps lightweight, task-specific forward/action heads:

INTACT inference and deployment pipeline

See Method Notes for the statistical construction, gradient contract, and the distinction from inverse dynamics, goal-conditioned behavior cloning, and post-hoc sampling priors.

Results

Task-Specific, One Epoch

The matched task-specific setting trains three models per task. Each checkpoint is evaluated with three seeds and 100 episodes per seed on the official LeWM protocol.

One-epoch direct control and local verification results

One epoch of goal-displacement INTACT reaches 95.61 +/- 0.59% Direct macro SR with no candidate search. The same checkpoints reach 96.58 +/- 0.44% with Guarded A (H=5, RH=5, 128x3, raw-action sigma=0.25, top-k 16). Its 384 sampled candidates are distinct from the one deterministic final-mean rescore. Published LeWM numbers use its separate 10-epoch CEM protocol and serve only as landscape context, not paired significance controls. Exact task values and the full matched inference matrix remain available in Audited Results.

One Shared Encoder

At epoch 5, Goal-displacement INTACT reaches 91.22 +/- 0.51% Direct macro SR with one encoder shared across all four visual domains. The matched shared LeWM baseline reaches 66.17 +/- 2.67% with CEM 300x30. With all INTACT action heads disabled, pure-CEM macro still rises to 70.08 +/- 1.13%, separating representation shaping from direct action readout.

Controlled shared-encoder success rates across four tasks

Theory Meets Measurement

At fixed state zz, INTACT treats two endpoint conditions as equivalent when they induce the same expert action law:

yzy    pE(az,y)=pE(az,y).y\sim_z y' \iff p_E^\star(a\mid z,y)=p_E^\star(a\mid z,y').

This conditional action quotient predicts that control should track the relation between predicted and expert action-law families, not task clustering or latent rank alone. Across 15 eligible goal-displacement checkpoints, predicted-expert kNN overlap correlates with Direct SR at r = 0.968 and linear CKA at r = 0.988; pointwise action R2R^2 reaches r = 0.983.

Action-family alignment diagnostics and their correlation with direct control

Community

Join the INTACT World Model Community for focused discussion around JEPA, LeWM, INTACT, representation learning, and efficient robot control. New results, reproductions, critiques, collaborations, and early ideas are all welcome.

Open the permanent Community page
The current WeChat invitation is maintained behind this stable link.

Citation

An arXiv preprint is available at arXiv:2607.26056. Please prefer the machine-readable CITATION.cff record.

@misc{sun2026intact,
  title         = {INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models},
  author        = {Sun, Junhan and Zhao, Hao and Zhang, Guofeng},
  year          = {2026},
  eprint        = {2607.26056},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2607.26056}
}

Acknowledgements

INTACT builds on the LeWM and stable-worldmodel research ecosystem. Evaluation is reported on the official LeWM protocol and, where explicitly marked, the separate CLEAR-LeWM evaluator. See NOTICE.md for provenance and licensing boundaries.

Questions: luoliibaqi4747@gmail.com