MMGR Benchmark

August 25, 2026 · View on GitHub

All numbers in this file are taken verbatim from the MMGR paper appendix.


Main results (Table 4 — 10 tasks × 10 models, final score)

Reconstruction note. The appendix digest does not reproduce a verbatim Table 4. It instead maps each Table 4 row to a source table: ARC → Table 12, Math → Table 21, the four Embodied Navigation rows → Table 26, and Physical Concepts / Sports → Table 36 (the two task-level rows, not the diagnostic "Average"). For Maze and Sudoku, the digest states the Table 4 row is "aggregated from the VLM-based overall metric" (Table 7 / Table 9) but gives no single aggregated scalar, so those cells are marked n/s. Cells for models that a task did not evaluate (Gemini text baselines outside Math; image models on Physical Commonsense) are also n/s. n/s = not specified in the paper appendix digest. All values are percentages.

TaskVeo-3Sora-2Wan-2.2Nano-bananaNano-banana ProGPT-4o-imageGPT-image-1.5Qwen-imageGemini-3-FlashGemini-3-Pro
T1 Mazen/sn/sn/sn/sn/sn/sn/sn/sn/sn/s
T2 Sudokun/sn/sn/sn/sn/sn/sn/sn/sn/sn/s
T3 ARC4.80%11.67%0.15%8.15%28.07%0.00%12.72%2.15%n/sn/s
T4 Math10.61%10.56%0.00%11.47%72.69%26.24%25.52%11.22%71.38%74.14%
T5 Last-Mile Nav.60.00%0.00%14.17%74.17%75.83%0.00%55.84%16.67%n/sn/s
T6 Top-down View Nav.19.49%3.39%5.09%11.11%33.05%3.39%26.27%5.08%n/sn/s
T7 3D R.-W. Nav.22.50%0.00%24.17%79.17%85.00%13.33%77.92%38.33%n/sn/s
T8 SLAG11.02%12.50%0.85%28.79%37.29%16.67%31.36%6.78%n/sn/s
T9 Physical Concepts41.67%76.00%26.67%n/sn/sn/sn/sn/sn/sn/s
T10 Sports60.00%64.00%21.33%n/sn/sn/sn/sn/sn/sn/s

T1 — Maze

Table 7: Quantitative results for the 2D Maze task using VLM-based evaluation method

Column headers: Maze Changed ↓ | Cross Wall ↓ | Target Achievement ↑ | Action Reflection ↑ | Overall ↑

Generator: Depth-First Search

ModelMaze Changed ↓Cross Wall ↓Target Achievement ↑Action Reflection ↑Overall ↑
Level: Easy (3×3–5×5)
Veo-315.50%25.50%60.50%1.00%42.00%
Sora-267.50%7.50%12.50%77.50%2.50%
Wan-2.235.00%79.17%10.00%7.50%1.67%
Nano-banana5.00%30.00%85.00%N/A15.50%
Nano-banana Pro5.00%42.50%90.00%N/A17.50%
GPT-4o-image95.00%5.00%72.50%N/A0.00%
GPT-image-1.53.33%13.33%90.00%N/A13.33%
Qwen-image5.00%23.33%65.00%N/A11.67%
Level: Medium (6×6–9×9)
Veo-30.50%25.64%50.78%0.00%38.72%
Sora-247.50%12.50%10.00%60.00%7.50%
Wan-2.210.83%90.83%23.33%28.33%1.67%
Nano-banana0.63%30.63%71.25%N/A4.38%
Nano-banana Pro0.00%25.00%82.50%N/A2.50%
GPT-4o-image72.50%10.00%82.50%N/A5.00%
GPT-image-1.50.83%5.83%80.83%N/A9.17%
Qwen-image2.50%28.33%44.17%N/A0.00%
Level: Hard (10×10–13×13)
Veo-30.00%18.50%60.00%1.50%51.50%
Sora-257.50%7.50%25.00%60.00%10.00%
Wan-2.26.67%80.83%20.00%35.00%5.00%
Nano-banana0.00%24.17%60.00%N/A0.83%
Nano-banana Pro0.00%12.50%80.00%N/A5.00%
GPT-4o-image62.50%5.00%77.50%N/A0.00%
GPT-image-1.52.50%9.17%86.67%N/A3.33%
Qwen-image7.50%11.67%42.50%N/A1.67%

Generator: Wilson's Algorithm

ModelMaze Changed ↓Cross Wall ↓Target Achievement ↑Action Reflection ↑Overall ↑
Level: Easy (3×3–5×5)
Veo-33.50%21.50%61.50%2.50%46.50%
Sora-267.50%10.00%15.00%40.00%5.00%
Wan-2.232.50%84.17%15.00%8.33%1.67%
Nano-banana10.50%47.50%81.00%N/A6.50%
Nano-banana Pro10.00%32.50%85.00%N/A12.50%
GPT-4o-image82.50%15.00%90.00%N/A2.50%
GPT-image-1.55.00%16.67%86.67%N/A5.00%
Qwen-image5.83%12.50%67.50%N/A20.00%
Level: Medium (6×6–9×9)
Veo-31.25%15.62%55.63%1.25%47.50%
Sora-262.50%12.50%20.00%47.50%10.00%
Wan-2.214.17%85.83%14.17%15.83%1.67%
Nano-banana0.00%36.88%70.00%N/A1.25%
Nano-banana Pro2.50%17.50%80.00%N/A2.50%
GPT-4o-image75.00%5.00%80.00%N/A5.00%
GPT-image-1.50.00%11.67%83.33%N/A4.17%
Qwen-image4.17%25.00%47.50%N/A0.00%
Level: Hard (10×10–13×13)
Veo-31.25%18.75%58.75%0.63%45.62%
Sora-245.00%10.00%10.00%52.50%2.50%
Wan-2.210.00%89.17%13.33%28.33%0.83%
Nano-banana1.25%27.50%60.62%N/A0.00%
Nano-banana Pro0.00%22.50%75.00%N/A5.00%
GPT-4o-image60.00%7.50%77.50%N/A7.50%
GPT-image-1.52.50%2.50%85.00%N/A4.17%
Qwen-image2.50%27.50%38.33%N/A0.00%

Table 8: Quantitative results for the 2D Maze task using pixel-based evaluation method

Generator: Depth-First Search

ModelMaze Changed ↓Cross Wall ↓Target Achievement ↑Action Reflection ↑Overall ↑
Level: Easy (3×3–5×5)
Veo-371.00%76.00%33.50%59.50%10.50%
Sora-2100.00%97.50%45.00%100.00%0.00%
Wan-2.22.50%90.00%25.00%87.50%0.00%
Nano-banana2.50%69.50%44.00%N/A14.00%
Nano-banana Pro7.50%42.50%30.00%N/A12.50%
GPT-4o-image100.00%77.50%0.00%N/A0.00%
GPT-image-1.596.67%87.50%8.33%N/A1.67%
Qwen-image60.00%89.17%95.83%N/A0.00%
Level: Medium (6×6–9×9)
Veo-3100.00%96.48%14.07%43.22%0.00%
Sora-2100.00%97.50%60.00%100.00%0.00%
Wan-2.27.50%100.00%15.00%82.50%0.00%
Nano-banana0.00%78.50%22.00%N/A2.00%
Nano-banana Pro0.00%30.00%37.50%N/A30.00%
GPT-4o-image100.00%77.50%0.00%N/A0.00%
GPT-image-1.547.50%88.33%20.00%N/A3.33%
Qwen-image82.50%84.17%98.33%N/A0.00%
Level: Hard (10×10–13×13)
Veo-3100.00%98.50%10.00%22.50%0.00%
Sora-2100.00%95.00%52.50%60.00%0.00%
Wan-2.20.00%97.50%5.00%67.50%0.00%
Nano-banana15.00%94.00%10.00%N/A1.00%
Nano-banana Pro0.00%50.00%7.50%N/A0.00%
GPT-4o-image100.00%62.50%0.00%N/A0.00%
GPT-image-1.51.67%90.00%5.83%N/A0.00%
Qwen-image98.33%88.33%100.00%N/A0.00%

Generator: Wilson's Algorithm

ModelMaze Changed ↓Cross Wall ↓Target Achievement ↑Action Reflection ↑Overall ↑
Level: Easy (3×3–5×5)
Veo-372.00%75.00%37.50%79.00%11.50%
Sora-2100.00%97.50%65.00%100.00%0.00%
Wan-2.20.00%82.50%17.50%90.00%0.00%
Nano-banana0.00%83.50%42.50%N/A5.00%
Nano-banana Pro0.00%65.00%30.00%N/A7.50%
GPT-4o-image100.00%92.50%10.00%N/A0.00%
GPT-image-1.598.33%94.17%20.00%N/A0.83%
Qwen-image49.17%90.83%89.17%N/A0.83%
Level: Medium (6×6–9×9)
Veo-3100.00%97.50%12.00%56.00%0.00%
Sora-2100.00%95.00%47.50%100.00%0.00%
Wan-2.25.00%100.00%15.00%95.00%0.00%
Nano-banana0.00%93.50%19.50%N/A0.00%
Nano-banana Pro2.50%82.50%20.00%N/A0.00%
GPT-4o-image100.00%82.50%0.00%N/A0.00%
GPT-image-1.545.00%97.50%14.17%N/A0.00%
Qwen-image80.83%85.83%99.17%N/A0.00%
Level: Hard (10×10–13×13)
Veo-3100.00%98.00%6.00%29.00%0.00%
Sora-2100.00%95.00%55.00%100.00%0.00%
Wan-2.210.00%100.00%0.00%92.50%0.00%
Nano-banana11.00%99.50%19.00%N/A0.00%
Nano-banana Pro0.00%82.50%15.00%N/A5.00%
GPT-4o-image100.00%67.50%0.00%N/A0.00%
GPT-image-1.50.00%97.50%5.00%N/A0.00%
Qwen-image100.00%88.33%100.00%N/A0.00%

Additional numeric findings (Failure Modes Analysis, C.5.3)

  • Average maze structure IoU: 0.3210 for GPT-4o-image; 0.1260 for Sora-2; 0.6573 for Veo-3.
  • Both Sora-2 and GPT-4o-image exhibit a 100% maze changing rate across all difficulty levels.
  • Veo-3 layout-failure ↔ wall-crossing co-occurrence: co-occur in 84.99% of all cases; 93.92% of layout failures involve wall crossing; 94.18% of wall crossing events trigger layout failure; Phi coefficient ϕ = 0.3821.
  • Figure 6 (Nano-banana image): 36,825 wall-crossing pixels (ratio = 0.19); path IoU = 0.34, precision = 0.53, recall = 0.49; target achievement = 0.70; layout match = 0.97, maze structure IoU = 0.92; path length ratio = 2.0, 13 loops; overall score = 0.067 (no gated), fails gated.
  • Figure 6 (Veo-3 video, 192/192 frames): wall-crossing ratio = 0.005; path IoU = 0.58, precision = 0.58, recall = 1.00; target achievement = 1.00; maze structure IoU = 0.80; path length ratio = 1.17, 13 loops; overall score = 0.34 (no gated), passes all gated criteria.

T2 — Sudoku

Table 9: Quantitative results for the Sudoku task using VLM-based evaluation method

Columns: Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑

Grid Size: 4×4 — Easy

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)72.40%98.40%37.22%92.00%1.20%
Sora-2 (video)100.00%95.92%25.78%32.65%0.00%
Wan-2.2 (video)100.00%100.00%17.22%4.67%0.00%
Nano-banana (image)18.61%17.20%81.14%N/A0.00%
Nano-banana Pro (image)35.25%1.33%97.05%N/A0.00%
GPT-4o-image (image)21.55%19.26%74.99%N/A0.00%
GPT-image-1.5 (image)63.33%32.06%64.57%N/A22.67%
Qwen-image (image)59.66%45.33%83.30%N/A0.00%

Grid Size: 4×4 — Medium

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)75.60%100.00%37.61%95.60%0.00%
Sora-2 (video)100.00%100.00%21.71%61.90%0.00%
Wan-2.2 (video)100.00%100.00%14.36%6.00%0.00%
Nano-banana (image)21.42%21.30%67.45%N/A0.00%
Nano-banana Pro (image)34.81%0.50%94.73%N/A0.00%
GPT-4o-image (image)33.16%22.45%55.54%N/A0.00%
GPT-image-1.5 (image)60.67%37.39%52.99%N/A10.67%
Qwen-image (image)37.82%43.67%82.83%N/A0.00%

Grid Size: 4×4 — Hard

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)70.00%100.00%30.38%94.80%0.00%
Sora-2 (video)100.00%91.30%20.03%60.87%0.00%
Wan-2.2 (video)99.33%100.00%14.79%6.00%0.00%
Nano-banana (image)25.84%25.80%58.79%N/A0.00%
Nano-banana Pro (image)36.53%0.83%91.08%N/A0.00%
GPT-4o-image (image)35.73%27.70%47.49%N/A0.00%
GPT-image-1.5 (image)81.33%39.11%40.39%N/A7.33%
Qwen-image (image)35.00%47.50%84.69%N/A0.00%

Grid Size: 9×9 — Easy

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)68.40%100.00%15.47%70.00%0.00%
Sora-2 (video)95.74%95.74%8.47%34.04%4.26%
Wan-2.2 (video)87.33%100.00%18.03%8.00%0.00%
Nano-banana (image)17.59%31.66%58.73%N/A0.00%
Nano-banana Pro (image)19.09%33.67%51.92%N/A0.00%
GPT-4o-image (image)69.39%59.57%20.23%N/A0.00%
GPT-image-1.5 (image)100.00%88.79%13.85%N/A0.00%
Qwen-image (image)29.14%5.78%72.02%N/A0.00%

Grid Size: 9×9 — Medium

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)66.00%99.20%13.66%72.00%0.40%
Sora-2 (video)100.00%100.00%9.43%30.00%0.00%
Wan-2.2 (video)85.33%100.00%16.13%2.00%0.00%
Nano-banana (image)21.87%35.54%50.43%N/A0.00%
Nano-banana Pro (image)14.91%28.40%46.47%N/A0.00%
GPT-4o-image (image)69.54%59.69%19.18%N/A0.00%
GPT-image-1.5 (image)100.00%89.14%14.62%N/A0.00%
Qwen-image (image)28.58%7.41%73.74%N/A0.00%

Grid Size: 9×9 — Hard

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)68.80%99.60%13.45%70.40%0.00%
Sora-2 (video)92.86%92.86%8.65%35.71%7.14%
Wan-2.2 (video)93.00%100.00%16.86%3.00%0.00%
Nano-banana (image)27.25%38.16%41.11%N/A0.00%
Nano-banana Pro (image)14.75%41.36%40.11%N/A0.00%
GPT-4o-image (image)69.80%57.66%15.76%N/A0.00%
GPT-image-1.5 (image)100.00%87.26%12.81%N/A0.00%
Qwen-image (image)27.26%10.44%73.59%N/A0.00%

Table 10: Quantitative results for the Sudoku task using OCR-based evaluation method

Columns: Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑

Grid Size: 4×4 — Easy

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)80.40%75.17%28.49%23.20%0.00%
Sora-2 (video)100.00%64.80%20.97%22.45%0.00%
Wan-2.2 (video)100.00%91.67%5.68%18.00%0.00%
Nano-banana (image)0.00%26.93%73.85%N/A32.80%
Nano-banana Pro (image)0.00%42.50%87.35%N/A14.00%
GPT-4o-image (image)0.00%50.15%57.63%N/A7.00%
GPT-image-1.5 (image)63.33%32.06%61.90%N/A22.67%
Qwen-image (image)97.33%95.72%3.84%N/A0.00%

Grid Size: 4×4 — Medium

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)86.00%90.50%25.47%35.60%0.00%
Sora-2 (video)100.00%59.52%25.20%33.33%0.00%
Wan-2.2 (video)100.00%93.17%7.49%10.00%0.00%
Nano-banana (image)0.00%35.97%53.96%N/A13.60%
Nano-banana Pro (image)0.00%43.33%85.62%N/A10.00%
GPT-4o-image (image)0.00%56.06%41.47%N/A0.75%
GPT-image-1.5 (image)60.67%37.39%50.99%N/A10.67%
Qwen-image (image)92.67%99.22%2.53%N/A0.00%

Grid Size: 4×4 — Hard

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)90.80%91.53%26.02%31.20%0.00%
Sora-2 (video)100.00%61.41%21.96%32.61%0.00%
Wan-2.2 (video)100.00%94.83%11.91%4.00%0.00%
Nano-banana (image)0.00%49.87%42.92%N/A5.20%
Nano-banana Pro (image)0.00%36.00%86.86%N/A18.00%
GPT-4o-image (image)0.00%58.29%36.80%N/A0.25%
GPT-image-1.5 (image)81.33%39.11%39.73%N/A7.33%
Qwen-image (image)91.33%99.28%2.92%N/A0.00%

Grid Size: 9×9 — Easy

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)88.00%99.91%5.75%92.40%0.00%
Sora-2 (video)100.00%99.84%6.63%55.32%0.00%
Wan-2.2 (video)100.00%100.00%3.72%96.00%0.00%
Nano-banana (image)0.00%96.07%15.70%N/A0.00%
Nano-banana Pro (image)0.00%55.22%45.68%N/A0.00%
GPT-4o-image (image)0.00%87.25%11.13%N/A0.00%
GPT-image-1.5 (image)100.00%88.79%13.18%N/A0.00%
Qwen-image (image)97.33%100.00%0.11%N/A0.00%

Grid Size: 9×9 — Medium

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)84.40%99.94%6.70%94.00%0.00%
Sora-2 (video)100.00%100.00%6.92%44.00%0.00%
Wan-2.2 (video)100.00%100.00%3.03%68.00%0.00%
Nano-banana (image)0.00%96.37%13.70%N/A0.00%
Nano-banana Pro (image)0.00%53.58%35.92%N/A0.00%
GPT-4o-image (image)0.00%85.93%10.47%N/A0.00%
GPT-image-1.5 (image)100.00%89.14%11.96%N/A0.00%
Qwen-image (image)92.67%100.00%0.11%N/A0.00%

Grid Size: 9×9 — Hard

ModelClues Changed ↓Constraints Violation ↓Completion Accuracy ↑Action Reflection ↑Overall ↑
Veo-3 (video)81.60%99.99%6.93%86.40%0.00%
Sora-2 (video)100.00%100.00%5.32%50.00%0.00%
Wan-2.2 (video)100.00%100.00%3.32%50.00%0.00%
Nano-banana (image)0.00%95.36%13.12%N/A0.00%
Nano-banana Pro (image)0.00%59.88%32.55%N/A0.00%
GPT-4o-image (image)0.00%86.25%9.79%N/A0.00%
GPT-image-1.5 (image)100.00%87.26%12.81%N/A0.00%
Qwen-image (image)91.33%100.00%0.08%N/A0.00%

Note (Figure 11 diagnostic examples, Veo-3): sudoku_easy_4x4_000.mp4 — 91.67% constraint violation, 0% completion accuracy, clues changed at [1,3] from 2→7 in frames 62–64; sudoku_easy_9x9_050.mp4 — 168 clue instances modified, 100% constraint violation, 4.44% completion accuracy, action reflection = 1.


T3 — ARC

Table 11: Distribution of 456 ARC cases across shape consistency and difficulty levels, separated by v1 and v2

Percentages indicate the proportion within each shape consistency group for the corresponding benchmark version.

VersionShape ConsistencyEasyMediumHardTotal
V1Match102 (38.8%)124 (47.1%)37 (14.1%)263
V1Mismatch34 (28.8%)57 (48.3%)27 (22.9%)118
V1Total13618164381
V2Match1 (1.9%)27 (50.9%)25 (47.2%)53
V2Mismatch/7 (31.8%)15 (68.2%)22
V2Total1344075
OverallMatch103 (31.0%)151 (45.4%)62 (23.6%)316
OverallMismatch34 (25.0%)64 (45.7%)42 (29.3%)140
OverallTotal137 (30.0%)215 (47.1%)104 (22.8%)456

Table 12: Final task-level results for the ARC task

Scores match the ARC row in Table 4; subsequent v1/v2 tables provide diagnostic split analyses rather than the final aggregation.

ModelTypeFinal score ↑
Veo-3Video4.80%
Sora-2Video11.67%
Wan-2.2Video0.15%
Nano-bananaImage8.15%
Nano-banana ProImage28.07%
GPT-4o-imageImage0.00%
GPT-image-1.5Image12.72%
Qwen-imageImage2.15%

Table 13: Diagnostic quantitative results for the ARC v1 split (381 cases)

(GPT-image-1.5 is not listed in Table 13.)

ModelPattern Recog. ↑Grid Integrity ↑Color Accuracy ↑Overall ↑
Video Models
Veo-317.32%32.98%8.22%5.16%
Sora-271.99%94.58%36.75%20.18%
Wan-2.20.61%13.04%0.17%0.17%
Image Models
Nano-banana28.42%55.79%12.63%9.21%
Nano-banana Pro61.98%84.73%40.42%30.54%
GPT-4o-image1.05%10.24%0.52%0.00%
Qwen-image1.31%4.46%0.52%0.52%

Table 14: Diagnostic quantitative results for the ARC v2 split (75 cases)

ModelPattern Recog. ↑Grid Integrity ↑Color Accuracy ↑Overall ↑
Video Models
Veo-317.78%31.11%6.22%4.00%
Sora-24.00%16.00%1.33%1.33%
Wan-2.20.00%5.78%0.00%0.00%
Image Models
Nano-banana18.67%42.67%8.00%2.67%
Nano-banana Pro62.50%83.93%44.64%30.36%
GPT-4o-image1.33%2.67%1.33%0.00%
Qwen-image1.33%5.33%1.33%1.33%

Table 15: Diagnostic quantitative breakdown results (Match and Mismatch) for the ARC v1 task (381 cases)

Model / CategoryPattern Recog. ↑Grid Integrity ↑Color Accuracy ↑Overall ↑
Video Models
Veo-3 — Match17.74%35.74%9.51%5.70%
Veo-3 — Mismatch16.38%26.84%5.37%3.95%
Sora-2 — Match67.29%93.46%35.05%17.76%
Sora-2 — Mismatch80.51%96.61%39.83%24.58%
Wan-2.2 — Match0.51%13.31%0.00%0.00%
Wan-2.2 — Mismatch0.85%12.43%0.56%0.56%
Image Models
Nano-banana — Match24.05%53.82%11.07%8.40%
Nano-banana — Mismatch38.14%60.17%16.10%11.02%
Nano-banana Pro — Match62.24%89.63%42.74%31.54%
Nano-banana Pro — Mismatch61.29%72.04%34.41%27.96%
GPT-4o-image — Match0.38%9.51%0.76%0.00%
GPT-4o-image — Mismatch2.54%11.86%0.00%0.00%
Qwen-image — Match0.76%4.56%0.76%0.76%
Qwen-image — Mismatch2.54%4.24%0.00%0.00%

Table 16: Diagnostic quantitative breakdown results for the ARC v1 task (381 cases) across difficulty levels (Easy, Medium, Hard)

Model / DifficultyPattern Recog. ↑Grid Integrity ↑Color Accuracy ↑Overall ↑
Veo-3: Match
Easy18.95%40.52%11.76%5.88%
Medium17.74%35.48%8.06%5.65%
Hard14.41%23.42%8.11%5.41%
Veo-3: Mismatch
Easy19.61%37.25%6.86%4.90%
Medium14.04%23.39%5.26%3.51%
Hard17.28%20.99%3.70%3.70%
Sora-2: Match
Easy73.17%90.24%42.68%19.51%
Medium63.81%95.24%33.33%19.05%
Hard62.96%96.30%18.52%7.41%
Sora-2: Mismatch
Easy85.29%91.18%41.18%29.41%
Medium80.70%100.00%36.84%19.30%
Hard74.07%96.30%44.44%29.63%
Wan-2.2: Match
Easy0.00%16.34%0.00%0.00%
Medium0.54%12.37%0.00%0.00%
Hard1.80%8.11%0.00%0.00%
Wan-2.2: Mismatch
Easy0.00%17.65%0.00%0.00%
Medium1.75%9.36%1.17%1.17%
Hard0.00%12.35%0.00%0.00%
Nano-banana: Match
Easy27.45%56.86%13.73%8.82%
Medium20.33%50.41%8.13%7.32%
Hard27.03%56.76%13.51%10.81%
Nano-banana: Mismatch
Easy47.06%85.29%11.76%11.76%
Medium35.09%52.63%19.30%10.53%
Hard33.33%44.44%14.81%11.11%
Nano-banana Pro: Match
Easy62.77%90.43%45.74%31.91%
Medium60.71%89.29%37.50%29.46%
Hard65.71%88.57%51.43%37.14%
Nano-banana Pro: Mismatch
Easy60.71%71.43%35.71%25.00%
Medium65.12%74.42%37.21%30.23%
Hard54.55%68.18%27.27%27.27%
GPT-4o-image: Match
Easy0.98%10.78%1.96%0.00%
Medium0.00%11.29%0.00%0.00%
Hard0.00%0.00%0.00%0.00%
GPT-4o-image: Mismatch
Easy0.00%11.76%0.00%0.00%
Medium1.75%10.53%0.00%0.00%
Hard7.41%14.81%0.00%0.00%
Qwen-image: Match
Easy0.98%5.88%0.98%0.98%
Medium0.00%3.23%0.00%0.00%
Hard2.70%5.41%2.70%2.70%
Qwen-image: Mismatch
Easy2.94%2.94%0.00%0.00%
Medium3.51%7.02%0.00%0.00%
Hard0.00%0.00%0.00%0.00%

Table 17: Diagnostic quantitative breakdown results (Match and Mismatch) for the ARC v2 task (75 cases)

Model / CategoryPattern Recog. ↑Grid Integrity ↑Color Accuracy ↑Overall ↑
Video Models
Veo-3 — Match17.61%32.08%6.92%3.77%
Veo-3 — Mismatch18.18%28.79%4.55%4.55%
Sora-2 — Match5.66%18.87%1.89%1.89%
Sora-2 — Mismatch0.00%9.09%0.00%0.00%
Wan-2.2 — Match0.00%4.40%0.00%0.00%
Wan-2.2 — Mismatch0.00%9.09%0.00%0.00%
Image Models
Nano-banana — Match18.87%45.28%11.32%3.77%
Nano-banana — Mismatch18.18%36.36%0.00%0.00%
Nano-banana Pro — Match64.29%88.10%50.00%33.33%
Nano-banana Pro — Mismatch57.14%71.43%28.57%21.43%
GPT-4o-image — Match0.00%1.89%0.00%0.00%
GPT-4o-image — Mismatch4.55%4.55%4.55%0.00%
Qwen-image — Match1.89%5.66%1.89%1.89%
Qwen-image — Mismatch0.00%4.55%0.00%0.00%

Table 18: Diagnostic quantitative breakdown results for the ARC v2 task (75 cases) across difficulty levels (Easy, Medium, Hard)

* The Easy level of ARC v2 contains only one evaluation case; therefore, a model solving this case correctly achieves a 100% score.

Model / DifficultyPattern Recog. ↑Grid Integrity ↑Color Accuracy ↑Overall ↑
Veo-3: Match
Easy0.00%33.33%0.00%0.00%
Medium20.99%34.57%8.64%6.17%
Hard14.67%29.33%5.33%1.33%
Veo-3: Mismatch
Medium23.81%42.86%4.76%4.76%
Hard15.56%22.22%4.44%4.44%
Sora-2: Match
Easy0.00%0.00%0.00%0.00%
Medium7.41%22.22%3.70%3.70%
Hard4.00%16.00%0.00%0.00%
Sora-2: Mismatch
Medium0.00%28.57%0.00%0.00%
Hard0.00%0.00%0.00%0.00%
Wan-2.2: Match
Easy0.00%33.33%0.00%0.00%
Medium0.00%4.94%0.00%0.00%
Hard0.00%2.67%0.00%0.00%
Wan-2.2: Mismatch
Medium0.00%9.52%0.00%0.00%
Hard0.00%8.89%0.00%0.00%
Nano-banana: Match
Easy100.00%100.00%100.00%100.00%*
Medium18.52%55.56%7.41%0.00%
Hard16.00%32.00%12.00%4.00%
Nano-banana: Mismatch
Medium42.86%57.14%0.00%0.00%
Hard6.67%26.67%0.00%0.00%
Nano-banana Pro: Match
Easy100.00%100.00%100.00%100.00%*
Medium66.67%95.24%61.90%38.10%
Hard60.00%80.00%35.00%25.00%
Nano-banana Pro: Mismatch
Medium80.00%100.00%40.00%20.00%
Hard44.44%55.56%22.22%22.22%
GPT-4o-image: Match
Easy0.00%0.00%0.00%0.00%
Medium0.00%3.70%0.00%0.00%
Hard0.00%0.00%0.00%0.00%
GPT-4o-image: Mismatch
Medium0.00%0.00%0.00%0.00%
Hard6.67%6.67%6.67%0.00%
Qwen-image: Match
Easy0.00%0.00%0.00%0.00%
Medium3.70%7.41%3.70%3.70%
Hard0.00%4.00%0.00%0.00%
Qwen-image: Mismatch
Medium0.00%0.00%0.00%0.00%
Hard0.00%6.67%0.00%0.00%

Additional narrative numbers (E.4)

  • Final task-level hierarchy: Nano-banana Pro 28.07%, GPT-image-1.5 12.72%, Sora-2 (strongest video) 11.67%.
  • Nano-banana Pro Overall — Easy 30.89%, Medium 30.39%, Hard 30.23%; Grid Integrity 86.18% → 86.74% → 77.91%. Sora-2 Overall — Easy 22.22%, Medium 16.33%, Hard 10.64%; Pattern Recognition 76.07% → 40.43%. Veo-3 ~5% across levels; Grid Integrity 39.66% → 32.40% → 24.04%.
  • v1→v2 collapse: Sora-2 20.18% → 1.33% (93% decline); Nano-banana 9.21% → 2.67% (71% decline); Nano-banana Pro 30.54% → 30.36% (stable). Veo-3 Grid Integrity 32.98% (v1) vs 31.11% (v2).
  • Metric cascade (Nano-banana v1): Grid Integrity 55.79% → Pattern Recognition 28.42% → Color Accuracy 12.63% → Overall 9.21%. Color Accuracy = 20–25% of Grid Integrity across all models.

T4 — Math

Table 19: Per-benchmark sample counts

DatasetLevelSample Count
GSM8KGrade School50
MATH500High School Competition50
AIME 2024Invitational Competition30
AIME 2025Invitational Competition30
Omni-MATHMulti-level (T0-T4)167
Total327

Table 20: Omni-MATH sample distribution across difficulty levels × categories

DifficultyAlgebraApplied MathCalculusDiscrete MathGeometryPrecalculusNumberOther
T0 (Easiest)45555541
T155454550
T255555551
T355555450
T4 (Hardest)55155450
Total242520252423242

(Omni-MATH category totals sum to 167.)

Table 21: Final task-level results for the Math task

ModelTypeFinal score ↑
Veo-3Video10.61%
Sora-2Video10.56%
Wan-2.2Video0.00%
Nano-bananaImage11.47%
Nano-banana ProImage72.69%
GPT-4o-imageImage26.24%
GPT-image-1.5Image25.52%
Qwen-imageImage11.22%
Gemini-3-FlashText baseline71.38%
Gemini-3-ProText baseline74.14%

Table 22: Diagnostic fine-grained results for the Math task across the original benchmark splits

Columns: Process Success Rate ↑ | Outcome Success Rate ↑ | Action Reflection ↑ | Overall Success Rate ↑ (Primary Metric). "–" = missing per-dataset result.

Dataset: GSM8K

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)12.00%74.00%12.00%12.00%
Sora-2 (Video)38.00%64.00%16.00%30.00%
Wan-2.2 (Video)2.00%2.00%0.00%2.00%
Nano-banana (Image)44.00%88.00%N/A42.00%
Nano-banana Pro (Image)97.83%97.83%N/A97.83%
GPT-4o-image (Image)80.00%83.48%N/A75.65%
Qwen-image (Image)44.00%44.00%N/A44.00%

Dataset: MATH500

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)20.00%52.00%14.00%18.00%
Sora-2 (Video)34.04%59.57%10.64%31.91%
Wan-2.2 (Video)3.33%6.00%0.00%3.33%
Nano-banana (Image)16.00%74.00%N/A16.00%
Nano-banana Pro (Image)91.84%91.84%N/A91.84%
GPT-4o-image (Image)29.14%42.45%N/A27.34%
Qwen-image (Image)N/A

Dataset: AIME24

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)8.33%8.33%5.00%1.67%
Sora-2 (Video)8.70%13.04%4.35%4.35%
Wan-2.2 (Video)15.56%22.22%11.11%15.56%
Nano-banana (Image)5.00%15.00%N/A0.00%
Nano-banana Pro (Image)63.64%36.36%N/A31.82%
GPT-4o-image (Image)3.33%10.00%N/A0.00%
Qwen-image (Image)1.12%1.12%N/A1.12%

Dataset: AIME25

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)3.33%11.67%1.67%3.33%
Sora-2 (Video)8.70%21.74%4.35%0.00%
Wan-2.2 (Video)5.56%10.00%3.33%5.56%
Nano-banana (Image)1.75%33.33%N/A1.75%
Nano-banana Pro (Image)71.43%90.48%N/A66.67%
GPT-4o-image (Image)0.00%3.33%N/A0.00%
Qwen-image (Image)4.00%4.00%N/A4.00%

Dataset: Omni-MATH

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)4.79%15.57%5.09%3.89%
Sora-2 (Video)0.62%1.88%6.88%0.62%
Wan-2.2 (Video)0.41%3.46%0.61%0.41%
Nano-banana (Image)3.90%39.94%N/A3.90%
Nano-banana Pro (Image)65.77%85.59%N/A63.06%
GPT-4o-image (Image)4.79%41.92%N/A4.79%
Qwen-image (Image)6.67%6.67%N/A6.67%

Table 23: Quantitative breakdown results for the Omni-MATH task (five difficulty levels T0–T4)

Columns: Process Success Rate ↑ | Outcome Success Rate ↑ | Action Reflection ↑ | Overall Success Rate ↑ (Primary Metric).

T0

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)6.06%12.12%3.03%6.06%
Sora-2 (Video)3.03%6.06%15.15%3.03%
Wan-2.2 (Video)0.98%4.90%2.94%0.98%
Nano-banana (Image)0.00%33.82%N/A0.00%
Nano-banana Pro (Image)53.85%84.62%N/A53.85%
GPT-4o-image (Image)0.00%0.00%N/A0.00%
Qwen-image (Image)22.00%22.00%N/A22.00%

T1

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)4.55%4.55%1.52%3.03%
Sora-2 (Video)0.00%3.12%3.12%0.00%
Wan-2.2 (Video)0.00%6.06%0.00%0.00%
Nano-banana (Image)0.00%19.70%N/A0.00%
Nano-banana Pro (Image)26.67%60.00%N/A26.67%
GPT-4o-image (Image)0.00%0.00%N/A0.00%
Qwen-image (Image)8.42%13.68%N/A8.42%

T2

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)5.56%16.67%2.78%5.56%
Sora-2 (Video)0.00%0.00%2.86%0.00%
Wan-2.2 (Video)0.93%3.70%0.00%0.93%
Nano-banana (Image)4.17%47.22%N/A4.17%
Nano-banana Pro (Image)70.00%96.67%N/A66.67%
GPT-4o-image (Image)2.78%2.78%N/A2.78%
Qwen-image (Image)7.92%8.91%N/A7.92%

T3

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)1.96%13.24%1.47%1.47%
Sora-2 (Video)0.00%0.00%9.68%0.00%
Wan-2.2 (Video)0.00%1.96%0.00%0.00%
Nano-banana (Image)5.88%45.59%N/A5.88%
Nano-banana Pro (Image)70.83%79.17%N/A62.50%
GPT-4o-image (Image)0.00%2.86%N/A0.00%
Qwen-image (Image)16.67%12.22%N/A12.22%

T4

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)5.17%36.67%15.00%5.00%
Sora-2 (Video)0.00%0.00%3.45%0.00%
Wan-2.2 (Video)0.00%0.00%0.00%0.00%
Nano-banana (Image)9.52%53.97%N/A9.52%
Nano-banana Pro (Image)82.76%93.10%N/A82.76%
GPT-4o-image (Image)0.00%6.67%N/A0.00%
Qwen-image (Image)25.68%22.97%N/A22.97%

Table 24: Quantitative breakdown results for the Omni-MATH task (eight categories)

Columns: Process Success Rate ↑ | Outcome Success Rate ↑ | Action Reflection ↑ | Overall Success Rate ↑ (Primary Metric).

Algebra

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)4.17%12.50%4.17%4.17%
Sora-2 (Video)4.35%8.70%13.04%4.35%
Wan-2.2 (Video)0.00%5.56%0.00%0.00%
Nano-banana (Image)4.17%45.83%N/A4.17%
Nano-banana Pro (Image)61.11%83.33%N/A55.56%
GPT-4o-image (Image)0.00%4.00%N/A0.00%
Qwen-image (Image)19.70%24.24%N/A19.70%

Applied Math

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)0.00%20.00%8.00%0.00%
Sora-2 (Video)0.00%0.00%8.00%0.00%
Wan-2.2 (Video)0.00%5.33%0.00%0.00%
Nano-banana (Image)16.00%32.00%N/A16.00%
Nano-banana Pro (Image)68.75%87.50%N/A68.75%
GPT-4o-image (Image)4.00%4.00%N/A4.00%
Qwen-image (Image)14.49%14.49%N/A14.49%

Calculus

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)5.00%10.00%5.00%0.00%
Sora-2 (Video)0.00%0.00%10.00%0.00%
Wan-2.2 (Video)0.00%1.67%0.00%0.00%
Nano-banana (Image)0.00%35.00%N/A0.00%
Nano-banana Pro (Image)66.67%80.00%N/A53.33%
GPT-4o-image (Image)0.00%0.00%N/A0.00%
Qwen-image (Image)15.79%15.79%N/A15.79%

Discrete Math

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)0.00%12.00%4.00%0.00%
Sora-2 (Video)0.00%0.00%4.17%0.00%
Wan-2.2 (Video)0.00%5.33%0.00%0.00%
Nano-banana (Image)0.00%24.00%N/A0.00%
Nano-banana Pro (Image)54.55%72.73%N/A54.55%
GPT-4o-image (Image)0.00%0.00%N/A0.00%
Qwen-image (Image)15.94%17.39%N/A15.94%

Geometry

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)8.33%25.00%4.17%8.33%
Sora-2 (Video)0.00%0.00%0.00%0.00%
Wan-2.2 (Video)0.00%1.39%4.17%0.00%
Nano-banana (Image)0.00%45.83%N/A0.00%
Nano-banana Pro (Image)78.57%92.86%N/A78.57%
GPT-4o-image (Image)0.00%0.00%N/A0.00%
Qwen-image (Image)23.64%20.00%N/A18.18%

Precalculus

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)4.35%21.74%8.70%0.00%
Sora-2 (Video)0.00%0.00%0.00%0.00%
Wan-2.2 (Video)0.00%1.45%0.00%0.00%
Nano-banana (Image)13.04%65.22%N/A13.04%
Nano-banana Pro (Image)68.18%86.36%N/A68.18%
GPT-4o-image (Image)0.00%4.17%N/A0.00%
Qwen-image (Image)20.90%16.42%N/A16.42%

Number

ModelProcess SROutcome SRAction ReflectionOverall SR
Veo-3 (Video)4.17%8.33%12.50%4.17%
Sora-2 (Video)0.00%0.00%0.00%0.00%
Wan-2.2 (Video)1.39%1.39%0.00%1.39%
Nano-banana (Image)0.00%50.00%N/A0.00%
Nano-banana Pro (Image)64.29%92.86%N/A64.29%
GPT-4o-image (Image)0.00%4.00%N/A0.00%
Qwen-image (Image)2.82%4.23%N/A2.82%

(Note: Table 24 in the source ends at the "Number" category; no "Other" category rows appear in the transcribed table.)


T5–T8 — Embodied Navigation (shared tables)

Table 25: Distribution of evaluation samples across the 24 hard-level configurations

Every one of the 24 configurations holds 5 samples for each of the four navigation tasks; each task totals 120.

Env. ComplexityView FidelityDistanceDestination TypeLast-Mile Nav.Top-down View Nav.3D R.-W. Nav.SLAG
1 Floorquality03shortcolor mark5555
1 Floorquality03shortlocation description5555
1 Floorquality03longcolor mark5555
1 Floorquality03longlocation description5555
1 Floorquality04shortcolor mark5555
1 Floorquality04shortlocation description5555
1 Floorquality04longcolor mark5555
1 Floorquality04longlocation description5555
1 Floorquality05shortcolor mark5555
1 Floorquality05shortlocation description5555
1 Floorquality05longcolor mark5555
1 Floorquality05longlocation description5555
2 Plus Floorsquality03shortcolor mark5555
2 Plus Floorsquality03shortlocation description5555
2 Plus Floorsquality03longcolor mark5555
2 Plus Floorsquality03longlocation description5555
2 Plus Floorsquality04shortcolor mark5555
2 Plus Floorsquality04shortlocation description5555
2 Plus Floorsquality04longcolor mark5555
2 Plus Floorsquality04longlocation description5555
2 Plus Floorsquality05shortcolor mark5555
2 Plus Floorsquality05shortlocation description5555
2 Plus Floorsquality05longcolor mark5555
2 Plus Floorsquality05longlocation description5555
Total (all 24 configs)120120120120

Table 26: Final task-level scores for Embodied Navigation (all values %)

TaskVeo-3Sora-2Wan-2.2Nano-bananaNano-banana ProGPT-4o-imageGPT-image-1.5Qwen-image
Last-Mile Nav.60.00%0.00%14.17%74.17%75.83%0.00%55.84%16.67%
Top-down View Nav.19.49%3.39%5.09%11.11%33.05%3.39%26.27%5.08%
3D R.-W. Nav.22.50%0.00%24.17%79.17%85.00%13.33%77.92%38.33%
SLAG11.02%12.50%0.85%28.79%37.29%16.67%31.36%6.78%

T5 — Panoramic View Last-Mile Navigation

Table 27: Quantitative results for the Panoramic View Last-Mile Navigation benchmark (VLM-based)

Columns: S.S.(3D) | O.S.(3D) | Obj. Sem. | Agent Con. | Spa. Ali. | Des. Inte. | Scene Con. | Succ(3D) Orig. Dest. | Physics Validness | Overall Success. Models: Veo-3, Sora-2 (video); Nano-banana, GPT-4o-image (image).

Group / ModelS.S.(3D)O.S.(3D)Obj. Sem.Agent Con.Spa. Ali.Des. Inte.Scene Con.Succ(3D) Orig. Dest.Physics ValidnessOverall Success
Env. Complexity — floor01
Veo-390.00%90.00%93.33%93.33%93.33%90.00%98.33%90.00%81.67%73.33%
Sora-20.00%1.67%93.33%88.33%93.33%0.00%0.00%0.00%81.67%0.00%
Nano-banana76.67%80.00%96.67%88.33%88.33%80.00%91.67%75.00%88.33%73.33%
GPT-4o-image0.00%0.00%100.00%1.67%0.00%0.00%0.00%0.00%0.00%0.00%
Env. Complexity — floor02plus
Veo-358.33%66.67%95.00%93.33%91.67%55.00%91.67%55.00%83.33%46.67%
Sora-20.00%0.00%81.36%77.97%76.27%0.00%1.69%0.00%64.41%0.00%
Nano-banana80.00%80.00%96.67%90.00%86.67%80.00%90.00%78.33%86.67%75.00%
GPT-4o-image0.00%0.00%91.67%0.00%0.00%0.00%0.00%0.00%0.00%0.00%
View Fidelity — quality03
Veo-372.50%80.00%92.50%92.50%90.00%70.00%92.50%70.00%77.50%55.00%
Sora-20.00%2.56%79.49%82.05%76.92%0.00%2.56%0.00%69.23%0.00%
Nano-banana62.50%62.50%92.50%87.50%82.50%62.50%80.00%62.50%82.50%60.00%
GPT-4o-image0.00%0.00%90.00%0.00%0.00%0.00%0.00%0.00%0.00%0.00%
View Fidelity — quality04
Veo-375.00%75.00%95.00%95.00%90.00%75.00%97.50%75.00%82.50%62.50%
Sora-20.00%0.00%85.00%80.00%87.50%0.00%0.00%0.00%67.50%0.00%
Nano-banana85.00%90.00%100.00%87.50%87.50%90.00%92.50%80.00%87.50%80.00%
GPT-4o-image0.00%0.00%100.00%0.00%0.00%0.00%0.00%0.00%0.00%0.00%
View Fidelity — quality05
Veo-375.00%80.00%95.00%92.50%97.50%72.50%95.00%72.50%87.50%62.50%
Sora-20.00%0.00%97.50%87.50%90.00%0.00%0.00%0.00%82.50%0.00%
Nano-banana87.50%87.50%97.50%92.50%92.50%87.50%100.00%87.50%92.50%82.50%
GPT-4o-image0.00%0.00%97.50%2.50%0.00%0.00%0.00%0.00%0.00%0.00%
Trajectory Distance — short
Veo-381.67%83.33%96.67%93.33%90.00%80.00%98.33%80.00%83.33%66.67%
Sora-20.00%0.00%88.33%83.33%85.00%0.00%1.67%0.00%76.67%0.00%
Nano-banana86.67%86.67%98.33%93.33%91.67%86.67%95.00%85.00%91.67%81.67%
GPT-4o-image0.00%0.00%95.00%1.67%0.00%0.00%0.00%0.00%0.00%0.00%
Trajectory Distance — long
Veo-366.67%73.33%91.67%93.33%95.00%65.00%91.67%65.00%81.67%53.33%
Sora-20.00%1.69%86.44%83.05%84.75%0.00%0.00%0.00%69.49%0.00%
Nano-banana70.00%73.33%95.00%85.00%83.33%73.33%86.67%68.33%83.33%66.67%
GPT-4o-image0.00%0.00%96.67%0.00%0.00%0.00%0.00%0.00%0.00%0.00%
Destination Spec. — color mark
Veo-383.33%88.33%96.67%93.33%93.33%80.00%93.33%80.00%85.00%70.00%
Sora-20.00%1.69%86.44%77.97%81.36%0.00%0.00%0.00%67.80%0.00%
Nano-banana76.67%80.00%95.00%85.00%85.00%81.67%80.00%86.67%73.33%81.67%
GPT-4o-image0.00%0.00%93.33%0.00%0.00%0.00%0.00%0.00%0.00%0.00%
Destination Spec. — location description
Veo-365.00%68.33%91.67%93.33%91.67%65.00%96.67%65.00%80.00%50.00%
Sora-20.00%0.00%88.33%88.33%88.33%0.00%1.67%0.00%78.33%0.00%
Nano-banana80.00%80.00%98.33%93.33%93.33%80.00%95.00%80.00%93.33%78.33%
GPT-4o-image0.00%0.00%98.33%1.67%0.00%0.00%0.00%0.00%0.00%0.00%

Source note (color-mark Nano-banana row): the raw text lists eleven values (76.67, 80.00, 95.00, 85.00, 85.00, 81.67, 80.00, 86.67, 73.33, 81.67, 70.00); the mapping above follows the ten-column header order and the trailing 70.00 is an ambiguous plain-text extraction artifact.

Table 28: Quantitative Geo-Align results for the Panoramic View Last-Mile Navigation task

Column groups: Task Success (SR, SPL); Path Fidelity (nDTW, sDTW); Trajectory Error (ATE, RPE); Path Length (Pred, GT); Navigation Error (Nav, Dir). Models: Veo-3, Sora-2, Wan-2.2. (Table 28 omits a quality04 view-fidelity level.)

Group / ModelSRSPLnDTWsDTWATERPEPredGTNavDir
Env. Complexity — floor01
Veo-383.3%50.6%53.9%48.7%0.910.376.803.560.6611.18
Sora-255.0%26.5%41.7%29.0%1.700.6011.433.561.4014.02
Wan-2.253.3%31.7%44.5%32.6%2.650.7716.483.561.4624.06
Env. Complexity — floor02plus
Veo-350.0%32.3%41.1%28.7%1.460.488.613.751.6019.47
Sora-225.0%9.9%32.0%13.7%1.850.5910.903.752.0017.13
Wan-2.233.9%14.9%34.5%17.7%1.740.6011.003.751.6116.61
View Fidelity — quality03
Veo-362.5%42.5%46.0%35.3%1.200.377.984.431.1012.39
Sora-222.5%10.7%32.2%11.8%2.090.5712.604.432.1416.71
Wan-2.237.5%22.5%38.3%21.5%3.410.7819.824.431.6817.28
View Fidelity — quality05
Veo-377.5%46.8%53.2%47.2%0.930.396.533.080.8314.28
Sora-250.0%23.0%40.8%28.1%1.640.6310.983.081.5914.80
Wan-2.252.5%29.2%45.0%31.8%1.410.539.013.081.2719.79
Trajectory Complexity — noturn
Veo-367.2%37.5%49.0%39.4%1.060.466.832.840.9114.98
Sora-248.3%21.4%40.8%26.9%1.500.579.062.841.4816.13
Wan-2.253.4%28.2%43.9%31.4%1.310.578.542.841.2518.55
Trajectory Complexity — oneturn
Veo-367.2%46.1%46.4%38.8%1.290.388.524.471.3215.38
Sora-232.8%15.6%33.3%16.4%2.040.6113.284.471.9114.92
Wan-2.234.5%18.9%35.5%19.4%3.110.8119.134.471.8122.38
Destination Spec. — color
Veo-363.8%39.3%47.9%37.5%1.080.417.403.651.1116.17
Sora-231.0%13.9%34.1%17.1%1.890.6611.933.651.8918.92
Wan-2.237.9%19.3%36.4%21.8%2.930.8517.643.651.8426.49
Destination Spec. — object
Veo-370.7%44.3%47.4%40.7%1.260.437.953.651.1214.19
Sora-250.0%23.1%39.9%26.2%1.650.5210.423.651.5012.12
Wan-2.250.0%27.9%43.0%29.0%1.490.5310.023.651.2214.44

H.3 narrative: auto-metrics rate Veo-3 at 73.33% success while humans rate it at only 25.00%. In floor01 Veo-3 and Nano-banana both reach 73.33% Overall Success; in floor02plus Veo-3 degrades to 46.67% while Nano-banana holds 75.00%. Nano-banana scales with fidelity 60.00% (quality03) → 82.50% (quality05); Veo-3 plateaus at ≤62.50%.


T6 — Top-down View Navigation

Table 29: Quantitative results for the 2D Top-down Navigation benchmark (VLM-based)

Compares Sora-2, Veo-3, Nano-banana, GPT-4o-image (all values %). Columns: Success Score (2D) | Oracle Success Score (2D) | Object Semantic | Agent Consistency | Spatial Alignment | Destination Integrity | Scene Consistency | Physics Validness (gate) | Overall Success (holistic).

Axis / LevelModelSuccess Score (2D)Oracle Success Score (2D)Object SemanticAgent ConsistencySpatial AlignmentDestination IntegrityScene ConsistencyPhysics ValidnessOverall Success
Env. Complexity — floor01Veo-365.7185.7177.1462.8691.4365.7145.7154.2937.14
Sora-222.8147.3782.4668.4278.957.028.7752.631.75
Nano-banana41.6752.7888.8950.0077.7833.3338.8950.005.56
GPT-4o-image10.3415.5255.176.9056.901.721.726.901.72
Env. Complexity — floor02plusVeo-325.0083.3361.1144.4480.5633.3345.7133.3313.89
Sora-228.8155.9389.8354.2486.4415.2523.7349.155.08
Nano-banana38.8952.7886.1155.5686.1136.1166.6750.0016.67
GPT-4o-image16.6723.3351.6711.6741.675.0010.0010.005.00
View Fidelity — quality03Veo-345.8379.1766.6733.3375.0050.0043.4829.1720.83
Sora-234.2152.6392.1163.1681.5818.4218.4257.897.89
Nano-banana41.6758.3383.3345.8375.0041.6758.3337.5012.50
GPT-4o-image21.0523.6865.7913.1657.895.265.2610.535.26
View Fidelity — quality04Veo-337.5083.3370.8358.3395.8341.6745.8354.1729.17
Sora-223.0846.1587.1866.6789.747.6917.9551.282.56
Nano-banana37.5050.0087.5050.0083.3333.3337.5050.004.17
GPT-4o-image12.5017.5052.5012.5050.005.0010.0012.505.00
View Fidelity — quality05Veo-352.1791.3069.5769.5786.9656.5247.8347.8326.09
Sora-220.5156.4179.4953.8576.927.6912.8243.590.00
Nano-banana41.6750.0091.6762.5087.5029.1762.5062.5016.67
GPT-4o-image7.5017.5042.502.5040.000.002.502.500.00
Trajectory Distance — noturnVeo-336.1183.3361.1152.7883.3338.8954.2938.8925.00
Sora-229.3160.3489.6662.0784.4815.5218.9748.281.72
Nano-banana41.6752.7888.8950.0083.3336.1152.7844.4411.11
GPT-4o-image20.3423.7357.6310.1757.635.088.4710.175.08
Trajectory Distance — oneturnVeo-354.2985.7177.1454.2988.5760.0037.1448.5725.71
Sora-222.4143.1082.7660.3481.036.9013.7953.455.17
Nano-banana38.8952.7886.1155.5680.5633.3352.7855.5611.11
GPT-4o-image6.7815.2549.158.4740.681.693.396.781.69
Destination Spec — color markVeo-352.7891.6780.5669.4491.6758.3363.8961.1138.89
Sora-217.2453.4577.5956.9077.5912.0724.1444.831.72
Nano-banana58.3377.7891.6744.4486.1155.5655.5641.6713.89
GPT-4o-image10.3413.7939.663.4534.481.723.453.451.72
Destination Spec — location descriptionVeo-337.1477.1457.1437.1480.0040.0026.4725.7111.43
Sora-234.4850.0094.8365.5287.9310.348.6256.905.17
Nano-banana22.2227.7883.3361.1177.7813.8950.0058.338.33
GPT-4o-image16.6725.0066.6715.0063.335.008.3313.335.00

Table 30: Top-down View Real-World Navigation — View-Syn (higher is better except LPIPS)

ModelModalityFVD Norm ↑FID Norm ↑SSIM ↑PSNR ↑LPIPS ↓
Sora-2Video0.70790.40727.210.8345
Veo-3Video0.67310.14716.650.8800
Wan-2.2Video0.62850.43337.810.7955
GPT-imageImage0.6153
Nano-bananaImage0.6153
Qwen-imageImage0.6003
GPT-image-1.5Image0.5247
Nano-banana-proImage0.4890

G.4 narrative: Top-down is the hardest embodied row overall. Nano-banana Pro leads at 33.05%, GPT-image-1.5 second at 26.27%, Veo-3 third at 19.49%.


T7 — 3D Real-World Navigation

Table 31: Quantitative results for the 3D Real-world Navigation benchmark (VLM-based)

Compares Sora-2, Veo-3, Nano-banana, GPT-4o-image (all values %). Columns: S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success.

Environmental Complexity — floor01

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)85.0091.6781.6786.6763.3353.3315.0011.6740.003.33
Sora-2 (Video)88.8997.2272.2283.3397.2277.7844.440.0055.000.00
Nano-banana (Image)75.0075.00100.0097.2297.2275.0086.1175.0097.2275.00
GPT-4o-image (Image)13.3313.3388.3348.3358.3313.3321.6713.3335.0013.33

Environmental Complexity — floor02plus

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)68.3378.3375.0068.3373.3338.3318.335.0033.330.00
Sora-2 (Video)72.2277.7869.4475.0083.3358.3336.110.0060.000.00
Nano-banana (Image)83.3386.1194.4497.2294.4486.11100.0083.3394.4483.33
GPT-4o-image (Image)15.0016.6775.0056.6751.6715.0015.0013.3340.0013.33

View Fidelity — quality03

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)65.0077.5072.5075.0070.0040.0015.005.0030.000.00
Sora-2 (Video)75.0083.3379.1783.3391.6766.6741.670.0067.500.00
Nano-banana (Image)75.0079.1795.8395.8391.6779.1791.6775.0091.6775.00
GPT-4o-image (Image)7.507.5080.0045.0052.505.007.505.0030.005.00

View Fidelity — quality04

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)82.5085.0075.0085.0065.0050.0020.0010.0042.502.50
Sora-2 (Video)91.6795.8354.1775.0083.3370.8341.670.0045.000.00
Nano-banana (Image)87.5087.50100.00100.00100.0087.5091.6787.50100.0087.50
GPT-4o-image (Image)17.5020.0080.0057.5052.5020.0027.5017.5040.0017.50

View Fidelity — quality05

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)82.5092.5087.5072.5070.0047.5015.0010.0037.502.50
Sora-2 (Video)75.0083.3379.1779.1795.8366.6737.500.0060.000.00
Nano-banana (Image)75.0075.0095.8395.8395.8375.0095.8375.0095.8375.00
GPT-4o-image (Image)17.5017.5085.0055.0060.0017.5020.0017.5042.5017.50

Trajectory Distance — short

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)73.3385.0081.6778.3368.3346.6716.678.3338.333.33
Sora-2 (Video)72.2283.3380.5677.7886.1163.8941.670.0055.000.00
Nano-banana (Image)77.7880.5697.2297.2294.4480.5688.8977.7894.4477.78
GPT-4o-image (Image)13.3313.3385.0056.6758.3311.6716.6711.6745.0011.67

Trajectory Distance — long

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)80.0085.0075.0076.6768.3345.0016.678.3335.000.00
Sora-2 (Video)88.8991.6761.1180.5694.4472.2238.890.0060.000.00
Nano-banana (Image)80.5680.5697.2297.2297.2280.5697.2280.5697.2280.56
GPT-4o-image (Image)15.0016.6778.3348.3351.6716.6720.0015.0030.0015.00

Destination Specification — color mark

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)93.3396.6766.6773.3370.0031.6728.3315.0028.333.33
Sora-2 (Video)88.8991.6758.3375.0088.8969.4436.110.0051.670.00
Nano-banana (Image)80.5680.56100.00100.00100.0080.5691.6780.56100.0080.56
GPT-4o-image (Image)16.6718.3371.6746.6745.0018.3320.0016.6730.0016.67

Destination Specification — location description

ModelS.S.(3D)O.S.(3D)Obj.Sem.Agent Con.Spa.Ali.Des.Inte.Scene Con.Success(3D) Orig.Dest.Physics ValidnessOverall Success
Veo-3 (Video)60.0073.3390.0081.6766.6760.005.001.6745.000.00
Sora-2 (Video)72.2283.3383.3383.3391.6766.6744.440.0063.330.00
Nano-banana (Image)77.7880.5694.4494.4491.6780.5694.4477.7891.6777.78
GPT-4o-image (Image)11.6711.6791.6758.3365.0010.0016.6710.0045.0010.00

Table 32: 3D Real-World Navigation — View-Syn (higher is better except LPIPS)

ModelModalityFVD Norm ↑FID Norm ↑SSIM ↑PSNR ↑LPIPS ↓
Veo-3Video0.77080.29877.350.8205
Wan-2.2Video0.71070.45629.110.7652
Sora-2Video0.68700.44719.190.8052
GPT-image-1.5Image0.6747
Nano-banana-proImage0.6740
GPT-imageImage0.5673
Nano-bananaImage0.5673
Qwen-imageImage0.5627

G.4.1: Nano-banana Pro reaches 85.00%, followed by Nano-banana (79.17%) and GPT-image-1.5 (77.92%). Video models remain below 25%.


T8 — SLAG

Table 33: Quantitative results for the SLAG benchmark (VLM-based)

Compares Sora-2, Veo-3, Nano-banana, GPT-4o-image (all values %). Column key: SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall.

Environmental Complexity — floor01

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-352.54%44.07%59.32%49.15%44.07%72.88%52.54%67.80%20.34%45.76%18.64%45.76%11.86%
Sora-229.31%27.59%34.48%37.93%55.17%84.48%84.48%72.41%32.76%79.31%22.41%65.52%10.34%
Nano-banana55.56%41.67%55.56%47.22%88.89%83.33%83.33%77.78%50.00%100.00%38.89%69.44%27.78%
GPT-4o-image35.00%28.33%36.67%31.67%66.67%81.67%76.67%61.67%26.67%98.33%25.00%58.33%25.00%

Environmental Complexity — floor02plus

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-321.05%38.60%26.32%42.11%33.33%50.88%42.11%52.63%31.58%57.89%15.79%31.58%10.53%
Sora-229.31%32.76%31.03%37.93%48.28%82.76%67.24%68.97%44.83%89.66%20.69%56.90%15.52%
Nano-banana40.00%56.67%40.00%63.33%76.67%86.67%86.67%70.00%50.00%96.67%36.67%60.00%30.00%
GPT-4o-image12.07%17.24%13.79%20.69%43.10%68.97%48.28%37.93%17.24%94.83%6.90%34.48%6.90%

View Fidelity — quality03

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-335.00%40.00%35.00%40.00%35.00%60.00%45.00%52.50%22.50%45.00%15.00%37.50%10.00%
Sora-227.50%27.50%27.50%32.50%47.50%85.00%77.50%75.00%37.50%77.50%22.50%62.50%17.50%
Nano-banana50.00%62.50%50.00%62.50%91.67%91.67%95.83%91.67%58.33%100.00%41.67%83.33%41.67%
GPT-4o-image30.00%25.00%32.50%32.50%65.00%75.00%65.00%52.50%20.00%97.50%17.50%52.50%17.50%

View Fidelity — quality04

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-346.15%51.28%51.28%58.97%43.59%66.67%46.15%69.23%33.33%48.72%25.64%35.90%15.38%
Sora-226.32%34.21%26.32%42.11%55.26%84.21%71.05%71.05%34.21%84.21%21.05%60.53%10.53%
Nano-banana45.83%45.83%45.83%54.17%70.83%87.50%79.17%58.33%50.00%95.83%37.50%54.17%20.83%
GPT-4o-image25.00%30.00%25.00%30.00%57.50%72.50%62.50%50.00%32.50%97.50%22.50%47.50%22.50%

View Fidelity — quality05

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-329.73%32.43%43.24%37.84%37.84%59.46%51.35%59.46%21.62%62.16%10.81%43.24%8.11%
Sora-234.21%28.95%44.74%39.47%52.63%81.58%78.95%65.79%44.74%92.11%21.05%60.53%10.53%
Nano-banana50.00%33.33%50.00%44.44%88.89%72.22%77.78%72.22%38.89%100.00%33.33%55.56%22.22%
GPT-4o-image15.79%13.16%18.42%15.79%42.11%78.95%60.53%47.37%13.16%94.74%7.89%39.47%7.89%

Trajectory Distance — short

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-338.98%40.68%45.76%45.76%44.07%61.02%54.24%64.41%28.81%62.71%20.34%42.37%15.25%
Sora-240.35%31.58%43.86%40.35%52.63%85.96%82.46%75.44%47.37%89.47%29.82%66.67%15.79%
Nano-banana54.55%51.52%54.55%54.55%87.88%90.91%90.91%90.91%57.58%100.00%42.42%78.79%36.36%
GPT-4o-image27.12%20.34%28.81%25.42%55.93%77.97%62.71%50.85%28.81%96.61%18.64%45.76%18.64%

Trajectory Distance — long

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-335.09%42.11%40.35%45.61%33.33%63.16%40.35%56.14%22.81%40.35%14.04%35.09%7.02%
Sora-218.64%28.81%22.03%35.59%50.85%81.36%69.49%66.10%30.51%79.66%13.56%55.93%10.17%
Nano-banana42.42%45.45%42.42%54.55%78.79%78.79%78.79%57.58%42.42%96.97%33.33%51.52%21.21%
GPT-4o-image20.34%25.42%22.03%27.12%54.24%72.88%62.71%49.15%15.25%96.61%13.56%47.46%13.56%

Destination Specification — color mark

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-353.45%51.72%60.34%56.90%43.10%56.90%46.55%65.52%32.76%67.24%25.86%34.48%17.24%
Sora-244.64%46.43%51.79%57.14%48.21%82.14%73.21%67.86%62.50%92.86%33.93%58.93%19.64%
Nano-banana81.25%59.38%81.25%68.75%87.50%84.38%87.50%78.12%71.88%100.00%62.50%71.88%50.00%
GPT-4o-image44.83%36.21%44.83%41.38%55.17%70.69%62.07%51.72%37.93%96.55%31.03%46.55%31.03%

Destination Specification — location description

ModelSS3DSS2DOS3DOS2DTrajAlignObjSemAgentConSpaAlignDestIntegSceneConSucc3D-OrigDestPhysValidOverall
Veo-320.69%31.03%25.86%34.48%34.48%67.24%48.28%55.17%18.97%36.21%8.62%43.10%5.17%
Sora-215.00%15.00%15.00%20.00%55.00%85.00%78.33%73.33%16.67%76.67%10.00%63.33%6.67%
Nano-banana17.65%38.24%17.65%41.18%79.41%85.29%82.35%70.59%29.41%97.06%14.71%58.82%8.82%
GPT-4o-image3.33%10.00%6.67%11.67%55.00%80.00%63.33%48.33%6.67%96.67%1.67%46.67%1.67%

Table 34 / K.3.2 View-Syn (SLAG)

The K.3.2 "View-Syn Evaluation" subsection header exists but contains no body text and no results table in the paper. View-Syn numeric results for SLAG: not specified in paper. (The digest also notes a Table 34 of FVD/FID/SSIM/PSNR/LPIPS is associated with SLAG but no numeric values were transcribed.)

G.4.1: Nano-banana Pro leads SLAG at 37.29%, followed by GPT-image-1.5 (31.36%) and Nano-banana (28.79%).


T9 & T10 — Physical Commonsense

Table 35: Distribution of Physical Commonsense evaluation samples

Task AxisTotal Samples
Physical Concepts (Atomic Interactions)25
Sports Scenarios (Compositional Contexts)25
Total50

Table 36: Quantitative results for the Physical Commonsense task

Models: Veo-3, Sora-2, Wan-2.2 (video generative models), evaluated with Gemini-2.5-Pro. Higher is better (↑) for all.

ModelPhysics Accuracy ↑Motion Quality ↑Visual Realism ↑Prompt Adherence ↑Overall ↑
Scenario Type: Physical Concepts
Veo-362.50%54.17%83.33%50.00%41.67%
Sora-284.00%80.00%96.00%76.00%76.00%
Wan-2.258.67%53.33%72.00%38.67%26.67%
Scenario Type: Sports Scenarios
Veo-380.00%68.00%92.00%68.00%60.00%
Sora-288.00%72.00%88.00%68.00%64.00%
Wan-2.242.67%33.33%96.00%21.33%21.33%
Average
Veo-371.43%61.22%87.76%59.18%51.02%
Sora-286.00%76.00%92.00%72.00%70.00%
Wan-2.250.67%43.33%84.00%30.00%24.00%

Note: the two task-level rows reported in Table 4 are Physical Concepts and Sports separately; the "Average" row here is a diagnostic aggregate.

Table 37: Fine-grained VLM-based evaluation results by Sports Scenarios attributes

Overall success rate (%) across sport types (Ballet, Diving, Skiing, Swimming) and difficulty levels (Easy, Medium, Hard).

ModelBalletDivingSkiingSwimmingEasyMediumHard
Veo-333.3%50.0%71.4%83.3%60.0%62.5%57.1%
Sora-233.3%50.0%85.7%83.3%60.0%75.0%57.1%
Wan-2.244.4%0.0%28.6%11.1%16.7%33.3%14.3%

Table 38: Fine-grained VLM-based evaluation results by Physical Concepts attributes

Overall success rate (%) across states-of-matter interaction types and difficulty levels.

ModelSolid-SolidSolid-FluidFluid-FluidAction/OtherEasyHard
Veo-30.0%75.0%50.0%40.0%53.3%22.2%
Sora-2100.0%75.0%100.0%66.7%75.0%77.8%
Wan-2.233.3%25.0%83.3%17.8%35.4%11.1%

Key findings (L.5): Sora-2 leads both final task rows — 76.00% on Physical Concepts and 64.00% on Sports (average 70.00%), ahead of Veo-3 (51.02%) and Wan-2.2 (24.00%). Wan-2.2 produces highly photorealistic videos (84–96% Visual Realism) yet fails prompts/physics (24% aggregate Overall). Solid-Solid interactions are most challenging (Veo-3 = 0% on rigid body collisions vs 75% on Solid-Fluid).


Evaluated models, sampling, totals, and cost

Evaluated models (exact names, versions, sizes)

Video generators (3):

  • Veo-3 — closed-source, API; parameter count not disclosed.
  • Sora-2 — closed-source, API; parameter count not disclosed.
  • Wan-2.2 — open-source; used in its image-to-video I2V-A14B Mixture-of-Experts configuration, 27B total parameters with about 14B active per denoising step.

Image generators (5):

  • Nano-banana — closed-source, API.
  • Nano-banana Pro — closed-source, API.
  • GPT-4o-image — closed-source, API.
  • GPT-image-1.5 — closed-source, API.
  • Qwen-image / Qwen-Image — open-source; 20B-parameter Multimodal Diffusion Transformer generating at 1024×1024 resolution.

LLM/VLM text-solution baselines (2, abstract-reasoning tasks only):

  • Gemini-3-Flash — closed-source, API.
  • Gemini-3-Pro — closed-source, API.
  • The Gemini text baselines produce generations only for Maze, Sudoku, and Math (Section A.1).

Judge / evaluator model (all tasks): Gemini-2.5-Pro (VLM judge for tasks without reliable pixel-/symbol-level verification). Deterministic evaluators: Maze pixel-based; Sudoku OCR-based (PaddleOCR).

Samples per prompt and totals

  • Five (5) outputs per prompt from each generative model (default settings for closed models, recommended settings for open ones; no task-specific tuning).
  • Total evaluation samples (task instances): 1,853.
  • Total generations (5 samples per instance):
    • A video model produces 9,265 generations.
    • An image model produces 9,015 generations (the eight non-physics tasks).
    • A Gemini text baseline produces 4,335 generations (Maze, Sudoku, and Math).

Compute (Section A)

  • On a single A100 80GB GPU: Qwen-Image renders one image in ≈40 s; Wan-2.2 produces one clip in ≈520 s.
  • Open-source GPU consumption: Qwen-Image ≈100 A100 GPU-hours; Wan-2.2 ≈1,340 A100 GPU-hours; ≈1,440 GPU-hours total.
  • Assumed rental rate: ≈$1.5 per A100 GPU-hour.
  • Assumed API prices: ≈$0.05 per image; $0.40 and $0.10 per second of generated video for Veo-3 and Sora-2 respectively; plus metered token usage for the Gemini baselines and the Gemini-2.5-Pro judge.
  • Deterministic Maze and Sudoku checks run on CPU at negligible cost.

Table 6: Per-model cost breakdown (Est. cost USD)

Model / groupEst. cost (USD)
Veo-3 (API)~30k
Sora-2 (API)~9k
Four closed image models (API) — Nano-banana, Nano-banana Pro, GPT-4o-image, GPT-image-1.5~1.8k
Gemini baselines + judge (API)~3k
Wan-2.2 (A100, ~1,340 h)~2.0k
Qwen-Image (A100, ~100 h)~0.15k
Total~46k

Estimated total inference cost of the benchmark: ~$46k, dominated by video generation.

URLs