Known Issues & Fixes

August 14, 2026 · View on GitHub

1. B200 / RTX 6000 Pro: Attention backend error

When running on NVIDIA B200 or RTX 6000 Pro GPUs with sglang, the default attention backend may fail. Add the following flags to your training command:

+actor_rollout_ref.rollout.engine_kwargs.sglang.attention_backend=flashinfer \
+actor_rollout_ref.rollout.engine_kwargs.sglang.mm_attention_backend=triton_attn

2. <think> is a reserved token on some families — do not use prompt_format: wm there

On Qwen2.5-VL and InternVL3, <think> is three ordinary text tokens. On Qwen3-VL, Qwen3.5 and GLM it is a single reserved control token tied to the model's own thinking channel, and the model will not emit it as text. Sokoban's wm format requires all four of <observation><think><answer><prediction>, and a response missing any of them has its whole action list discarded (strict_format, deliberate). So on those families wm scores zero while every other metric looks healthy.

Measured on Qwen3-VL-4B: wm → format 0.000 / score 0.000; free_wm → format 0.969 / score 0.602.

Use free_wm (observation/answer/prediction, free prose between) or free_think (</think> then <answer>) on those families.

Sokoban and primitive_skill both default to wm, so a <think>-reserving model needs prompt_format set explicitly in the dataset yaml on either. frozenlake and navigation default to free_think and are unaffected. What each environment offers:

envformatsdefault
sokobanwm, free_wm, free_think, answerwm
primitive_skillwm, free_thinkwm
frozenlakewm, free_thinkfree_think
navigationwm, free_think, no_think, eval_modefree_think

free_wm and answer exist only for sokoban.

3. thinking_token_budget bounds the think block, not the response — and not the cost

For a model with a native reasoning channel (Qwen3-VL, Qwen3.5, GLM), thinking_token_budget is passed to vLLM, which forces the closing </think> once the budget is spent. Measured against a live vllm serve Qwen/Qwen3.5-4B with --reasoning-config:

thinking_token_budgetthink tokensfinish_reasontotal tokens
unset889length (hit max_tokens)2048
512511stop641
128127stop1413

The parameter is exact. But vLLM forces the closing token mid-word, and the model does not treat that as having finished reasoning — it carries on in the same register inside the visible content, often never reaching <answer>. A tighter think budget can produce a longer response: 1413 total tokens at budget 128 against 641 at budget 512, because what it still wanted to say moved past the tag.

So the budget buys a well-formed </think> and a scoreable turn, not brevity. For brevity, change the prompt format — free_think is the only sokoban format whose instructions tell the model to stop reasoning, and it is what the shipped Qwen3.5 thinking config uses.

vLLM refuses the request if the budget is set and reasoning_config is not; see examples/train/sokoban/train_default_gae_qwen35_4b_think.sh for the delimiters, which are per-model-family knowledge and so live in the script rather than in VAGEN.

4. A compact_budget too small fails at runtime, not at startup

Under trainer.harness=compact, a budget that cannot hold the system prompt plus a summary plus one generation closes every conversation after a single turn, raises CompactionMakesNoProgress on every episode, and empties the batch. Measured at compact_budget=400 with a 589-token sokoban system prompt.

There is no static check, and there deliberately isn't one: the threshold depends on the system prompt, which vagen/harness/budget.py cannot see, and a version of the check that guessed refused three configurations that demonstrably run. The runtime error names this as the likely cause, but the symptom arrives after the allocation is up.

If a compact run empties its batch immediately, suspect this number first. 1200 is what train_default_gae_compact_qwen25vl3b.sh uses at max_turns: 5; the shipped default of 4000 holds a whole 5-turn sokoban episode, so nothing is ever summarised and the run is concat under another name.