MMGR Benchmark
August 25, 2026 · View on GitHub
All numbers in this file are taken verbatim from the MMGR paper appendix.
Main results (Table 4 — 10 tasks × 10 models, final score)
Reconstruction note. The appendix digest does not reproduce a verbatim Table 4. It instead maps each Table 4 row to a source table: ARC → Table 12, Math → Table 21, the four Embodied Navigation rows → Table 26, and Physical Concepts / Sports → Table 36 (the two task-level rows, not the diagnostic "Average"). For Maze and Sudoku, the digest states the Table 4 row is "aggregated from the VLM-based overall metric" (Table 7 / Table 9) but gives no single aggregated scalar, so those cells are marked
n/s. Cells for models that a task did not evaluate (Gemini text baselines outside Math; image models on Physical Commonsense) are alson/s.n/s= not specified in the paper appendix digest. All values are percentages.
| Task | Veo-3 | Sora-2 | Wan-2.2 | Nano-banana | Nano-banana Pro | GPT-4o-image | GPT-image-1.5 | Qwen-image | Gemini-3-Flash | Gemini-3-Pro |
|---|---|---|---|---|---|---|---|---|---|---|
| T1 Maze | n/s | n/s | n/s | n/s | n/s | n/s | n/s | n/s | n/s | n/s |
| T2 Sudoku | n/s | n/s | n/s | n/s | n/s | n/s | n/s | n/s | n/s | n/s |
| T3 ARC | 4.80% | 11.67% | 0.15% | 8.15% | 28.07% | 0.00% | 12.72% | 2.15% | n/s | n/s |
| T4 Math | 10.61% | 10.56% | 0.00% | 11.47% | 72.69% | 26.24% | 25.52% | 11.22% | 71.38% | 74.14% |
| T5 Last-Mile Nav. | 60.00% | 0.00% | 14.17% | 74.17% | 75.83% | 0.00% | 55.84% | 16.67% | n/s | n/s |
| T6 Top-down View Nav. | 19.49% | 3.39% | 5.09% | 11.11% | 33.05% | 3.39% | 26.27% | 5.08% | n/s | n/s |
| T7 3D R.-W. Nav. | 22.50% | 0.00% | 24.17% | 79.17% | 85.00% | 13.33% | 77.92% | 38.33% | n/s | n/s |
| T8 SLAG | 11.02% | 12.50% | 0.85% | 28.79% | 37.29% | 16.67% | 31.36% | 6.78% | n/s | n/s |
| T9 Physical Concepts | 41.67% | 76.00% | 26.67% | n/s | n/s | n/s | n/s | n/s | n/s | n/s |
| T10 Sports | 60.00% | 64.00% | 21.33% | n/s | n/s | n/s | n/s | n/s | n/s | n/s |
T1 — Maze
Table 7: Quantitative results for the 2D Maze task using VLM-based evaluation method
Column headers: Maze Changed ↓ | Cross Wall ↓ | Target Achievement ↑ | Action Reflection ↑ | Overall ↑
Generator: Depth-First Search
| Model | Maze Changed ↓ | Cross Wall ↓ | Target Achievement ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Level: Easy (3×3–5×5) | |||||
| Veo-3 | 15.50% | 25.50% | 60.50% | 1.00% | 42.00% |
| Sora-2 | 67.50% | 7.50% | 12.50% | 77.50% | 2.50% |
| Wan-2.2 | 35.00% | 79.17% | 10.00% | 7.50% | 1.67% |
| Nano-banana | 5.00% | 30.00% | 85.00% | N/A | 15.50% |
| Nano-banana Pro | 5.00% | 42.50% | 90.00% | N/A | 17.50% |
| GPT-4o-image | 95.00% | 5.00% | 72.50% | N/A | 0.00% |
| GPT-image-1.5 | 3.33% | 13.33% | 90.00% | N/A | 13.33% |
| Qwen-image | 5.00% | 23.33% | 65.00% | N/A | 11.67% |
| Level: Medium (6×6–9×9) | |||||
| Veo-3 | 0.50% | 25.64% | 50.78% | 0.00% | 38.72% |
| Sora-2 | 47.50% | 12.50% | 10.00% | 60.00% | 7.50% |
| Wan-2.2 | 10.83% | 90.83% | 23.33% | 28.33% | 1.67% |
| Nano-banana | 0.63% | 30.63% | 71.25% | N/A | 4.38% |
| Nano-banana Pro | 0.00% | 25.00% | 82.50% | N/A | 2.50% |
| GPT-4o-image | 72.50% | 10.00% | 82.50% | N/A | 5.00% |
| GPT-image-1.5 | 0.83% | 5.83% | 80.83% | N/A | 9.17% |
| Qwen-image | 2.50% | 28.33% | 44.17% | N/A | 0.00% |
| Level: Hard (10×10–13×13) | |||||
| Veo-3 | 0.00% | 18.50% | 60.00% | 1.50% | 51.50% |
| Sora-2 | 57.50% | 7.50% | 25.00% | 60.00% | 10.00% |
| Wan-2.2 | 6.67% | 80.83% | 20.00% | 35.00% | 5.00% |
| Nano-banana | 0.00% | 24.17% | 60.00% | N/A | 0.83% |
| Nano-banana Pro | 0.00% | 12.50% | 80.00% | N/A | 5.00% |
| GPT-4o-image | 62.50% | 5.00% | 77.50% | N/A | 0.00% |
| GPT-image-1.5 | 2.50% | 9.17% | 86.67% | N/A | 3.33% |
| Qwen-image | 7.50% | 11.67% | 42.50% | N/A | 1.67% |
Generator: Wilson's Algorithm
| Model | Maze Changed ↓ | Cross Wall ↓ | Target Achievement ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Level: Easy (3×3–5×5) | |||||
| Veo-3 | 3.50% | 21.50% | 61.50% | 2.50% | 46.50% |
| Sora-2 | 67.50% | 10.00% | 15.00% | 40.00% | 5.00% |
| Wan-2.2 | 32.50% | 84.17% | 15.00% | 8.33% | 1.67% |
| Nano-banana | 10.50% | 47.50% | 81.00% | N/A | 6.50% |
| Nano-banana Pro | 10.00% | 32.50% | 85.00% | N/A | 12.50% |
| GPT-4o-image | 82.50% | 15.00% | 90.00% | N/A | 2.50% |
| GPT-image-1.5 | 5.00% | 16.67% | 86.67% | N/A | 5.00% |
| Qwen-image | 5.83% | 12.50% | 67.50% | N/A | 20.00% |
| Level: Medium (6×6–9×9) | |||||
| Veo-3 | 1.25% | 15.62% | 55.63% | 1.25% | 47.50% |
| Sora-2 | 62.50% | 12.50% | 20.00% | 47.50% | 10.00% |
| Wan-2.2 | 14.17% | 85.83% | 14.17% | 15.83% | 1.67% |
| Nano-banana | 0.00% | 36.88% | 70.00% | N/A | 1.25% |
| Nano-banana Pro | 2.50% | 17.50% | 80.00% | N/A | 2.50% |
| GPT-4o-image | 75.00% | 5.00% | 80.00% | N/A | 5.00% |
| GPT-image-1.5 | 0.00% | 11.67% | 83.33% | N/A | 4.17% |
| Qwen-image | 4.17% | 25.00% | 47.50% | N/A | 0.00% |
| Level: Hard (10×10–13×13) | |||||
| Veo-3 | 1.25% | 18.75% | 58.75% | 0.63% | 45.62% |
| Sora-2 | 45.00% | 10.00% | 10.00% | 52.50% | 2.50% |
| Wan-2.2 | 10.00% | 89.17% | 13.33% | 28.33% | 0.83% |
| Nano-banana | 1.25% | 27.50% | 60.62% | N/A | 0.00% |
| Nano-banana Pro | 0.00% | 22.50% | 75.00% | N/A | 5.00% |
| GPT-4o-image | 60.00% | 7.50% | 77.50% | N/A | 7.50% |
| GPT-image-1.5 | 2.50% | 2.50% | 85.00% | N/A | 4.17% |
| Qwen-image | 2.50% | 27.50% | 38.33% | N/A | 0.00% |
Table 8: Quantitative results for the 2D Maze task using pixel-based evaluation method
Generator: Depth-First Search
| Model | Maze Changed ↓ | Cross Wall ↓ | Target Achievement ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Level: Easy (3×3–5×5) | |||||
| Veo-3 | 71.00% | 76.00% | 33.50% | 59.50% | 10.50% |
| Sora-2 | 100.00% | 97.50% | 45.00% | 100.00% | 0.00% |
| Wan-2.2 | 2.50% | 90.00% | 25.00% | 87.50% | 0.00% |
| Nano-banana | 2.50% | 69.50% | 44.00% | N/A | 14.00% |
| Nano-banana Pro | 7.50% | 42.50% | 30.00% | N/A | 12.50% |
| GPT-4o-image | 100.00% | 77.50% | 0.00% | N/A | 0.00% |
| GPT-image-1.5 | 96.67% | 87.50% | 8.33% | N/A | 1.67% |
| Qwen-image | 60.00% | 89.17% | 95.83% | N/A | 0.00% |
| Level: Medium (6×6–9×9) | |||||
| Veo-3 | 100.00% | 96.48% | 14.07% | 43.22% | 0.00% |
| Sora-2 | 100.00% | 97.50% | 60.00% | 100.00% | 0.00% |
| Wan-2.2 | 7.50% | 100.00% | 15.00% | 82.50% | 0.00% |
| Nano-banana | 0.00% | 78.50% | 22.00% | N/A | 2.00% |
| Nano-banana Pro | 0.00% | 30.00% | 37.50% | N/A | 30.00% |
| GPT-4o-image | 100.00% | 77.50% | 0.00% | N/A | 0.00% |
| GPT-image-1.5 | 47.50% | 88.33% | 20.00% | N/A | 3.33% |
| Qwen-image | 82.50% | 84.17% | 98.33% | N/A | 0.00% |
| Level: Hard (10×10–13×13) | |||||
| Veo-3 | 100.00% | 98.50% | 10.00% | 22.50% | 0.00% |
| Sora-2 | 100.00% | 95.00% | 52.50% | 60.00% | 0.00% |
| Wan-2.2 | 0.00% | 97.50% | 5.00% | 67.50% | 0.00% |
| Nano-banana | 15.00% | 94.00% | 10.00% | N/A | 1.00% |
| Nano-banana Pro | 0.00% | 50.00% | 7.50% | N/A | 0.00% |
| GPT-4o-image | 100.00% | 62.50% | 0.00% | N/A | 0.00% |
| GPT-image-1.5 | 1.67% | 90.00% | 5.83% | N/A | 0.00% |
| Qwen-image | 98.33% | 88.33% | 100.00% | N/A | 0.00% |
Generator: Wilson's Algorithm
| Model | Maze Changed ↓ | Cross Wall ↓ | Target Achievement ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Level: Easy (3×3–5×5) | |||||
| Veo-3 | 72.00% | 75.00% | 37.50% | 79.00% | 11.50% |
| Sora-2 | 100.00% | 97.50% | 65.00% | 100.00% | 0.00% |
| Wan-2.2 | 0.00% | 82.50% | 17.50% | 90.00% | 0.00% |
| Nano-banana | 0.00% | 83.50% | 42.50% | N/A | 5.00% |
| Nano-banana Pro | 0.00% | 65.00% | 30.00% | N/A | 7.50% |
| GPT-4o-image | 100.00% | 92.50% | 10.00% | N/A | 0.00% |
| GPT-image-1.5 | 98.33% | 94.17% | 20.00% | N/A | 0.83% |
| Qwen-image | 49.17% | 90.83% | 89.17% | N/A | 0.83% |
| Level: Medium (6×6–9×9) | |||||
| Veo-3 | 100.00% | 97.50% | 12.00% | 56.00% | 0.00% |
| Sora-2 | 100.00% | 95.00% | 47.50% | 100.00% | 0.00% |
| Wan-2.2 | 5.00% | 100.00% | 15.00% | 95.00% | 0.00% |
| Nano-banana | 0.00% | 93.50% | 19.50% | N/A | 0.00% |
| Nano-banana Pro | 2.50% | 82.50% | 20.00% | N/A | 0.00% |
| GPT-4o-image | 100.00% | 82.50% | 0.00% | N/A | 0.00% |
| GPT-image-1.5 | 45.00% | 97.50% | 14.17% | N/A | 0.00% |
| Qwen-image | 80.83% | 85.83% | 99.17% | N/A | 0.00% |
| Level: Hard (10×10–13×13) | |||||
| Veo-3 | 100.00% | 98.00% | 6.00% | 29.00% | 0.00% |
| Sora-2 | 100.00% | 95.00% | 55.00% | 100.00% | 0.00% |
| Wan-2.2 | 10.00% | 100.00% | 0.00% | 92.50% | 0.00% |
| Nano-banana | 11.00% | 99.50% | 19.00% | N/A | 0.00% |
| Nano-banana Pro | 0.00% | 82.50% | 15.00% | N/A | 5.00% |
| GPT-4o-image | 100.00% | 67.50% | 0.00% | N/A | 0.00% |
| GPT-image-1.5 | 0.00% | 97.50% | 5.00% | N/A | 0.00% |
| Qwen-image | 100.00% | 88.33% | 100.00% | N/A | 0.00% |
Additional numeric findings (Failure Modes Analysis, C.5.3)
- Average maze structure IoU: 0.3210 for GPT-4o-image; 0.1260 for Sora-2; 0.6573 for Veo-3.
- Both Sora-2 and GPT-4o-image exhibit a 100% maze changing rate across all difficulty levels.
- Veo-3 layout-failure ↔ wall-crossing co-occurrence: co-occur in 84.99% of all cases; 93.92% of layout failures involve wall crossing; 94.18% of wall crossing events trigger layout failure; Phi coefficient ϕ = 0.3821.
- Figure 6 (Nano-banana image): 36,825 wall-crossing pixels (ratio = 0.19); path IoU = 0.34, precision = 0.53, recall = 0.49; target achievement = 0.70; layout match = 0.97, maze structure IoU = 0.92; path length ratio = 2.0, 13 loops; overall score = 0.067 (no gated), fails gated.
- Figure 6 (Veo-3 video, 192/192 frames): wall-crossing ratio = 0.005; path IoU = 0.58, precision = 0.58, recall = 1.00; target achievement = 1.00; maze structure IoU = 0.80; path length ratio = 1.17, 13 loops; overall score = 0.34 (no gated), passes all gated criteria.
T2 — Sudoku
Table 9: Quantitative results for the Sudoku task using VLM-based evaluation method
Columns: Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑
Grid Size: 4×4 — Easy
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 72.40% | 98.40% | 37.22% | 92.00% | 1.20% |
| Sora-2 (video) | 100.00% | 95.92% | 25.78% | 32.65% | 0.00% |
| Wan-2.2 (video) | 100.00% | 100.00% | 17.22% | 4.67% | 0.00% |
| Nano-banana (image) | 18.61% | 17.20% | 81.14% | N/A | 0.00% |
| Nano-banana Pro (image) | 35.25% | 1.33% | 97.05% | N/A | 0.00% |
| GPT-4o-image (image) | 21.55% | 19.26% | 74.99% | N/A | 0.00% |
| GPT-image-1.5 (image) | 63.33% | 32.06% | 64.57% | N/A | 22.67% |
| Qwen-image (image) | 59.66% | 45.33% | 83.30% | N/A | 0.00% |
Grid Size: 4×4 — Medium
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 75.60% | 100.00% | 37.61% | 95.60% | 0.00% |
| Sora-2 (video) | 100.00% | 100.00% | 21.71% | 61.90% | 0.00% |
| Wan-2.2 (video) | 100.00% | 100.00% | 14.36% | 6.00% | 0.00% |
| Nano-banana (image) | 21.42% | 21.30% | 67.45% | N/A | 0.00% |
| Nano-banana Pro (image) | 34.81% | 0.50% | 94.73% | N/A | 0.00% |
| GPT-4o-image (image) | 33.16% | 22.45% | 55.54% | N/A | 0.00% |
| GPT-image-1.5 (image) | 60.67% | 37.39% | 52.99% | N/A | 10.67% |
| Qwen-image (image) | 37.82% | 43.67% | 82.83% | N/A | 0.00% |
Grid Size: 4×4 — Hard
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 70.00% | 100.00% | 30.38% | 94.80% | 0.00% |
| Sora-2 (video) | 100.00% | 91.30% | 20.03% | 60.87% | 0.00% |
| Wan-2.2 (video) | 99.33% | 100.00% | 14.79% | 6.00% | 0.00% |
| Nano-banana (image) | 25.84% | 25.80% | 58.79% | N/A | 0.00% |
| Nano-banana Pro (image) | 36.53% | 0.83% | 91.08% | N/A | 0.00% |
| GPT-4o-image (image) | 35.73% | 27.70% | 47.49% | N/A | 0.00% |
| GPT-image-1.5 (image) | 81.33% | 39.11% | 40.39% | N/A | 7.33% |
| Qwen-image (image) | 35.00% | 47.50% | 84.69% | N/A | 0.00% |
Grid Size: 9×9 — Easy
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 68.40% | 100.00% | 15.47% | 70.00% | 0.00% |
| Sora-2 (video) | 95.74% | 95.74% | 8.47% | 34.04% | 4.26% |
| Wan-2.2 (video) | 87.33% | 100.00% | 18.03% | 8.00% | 0.00% |
| Nano-banana (image) | 17.59% | 31.66% | 58.73% | N/A | 0.00% |
| Nano-banana Pro (image) | 19.09% | 33.67% | 51.92% | N/A | 0.00% |
| GPT-4o-image (image) | 69.39% | 59.57% | 20.23% | N/A | 0.00% |
| GPT-image-1.5 (image) | 100.00% | 88.79% | 13.85% | N/A | 0.00% |
| Qwen-image (image) | 29.14% | 5.78% | 72.02% | N/A | 0.00% |
Grid Size: 9×9 — Medium
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 66.00% | 99.20% | 13.66% | 72.00% | 0.40% |
| Sora-2 (video) | 100.00% | 100.00% | 9.43% | 30.00% | 0.00% |
| Wan-2.2 (video) | 85.33% | 100.00% | 16.13% | 2.00% | 0.00% |
| Nano-banana (image) | 21.87% | 35.54% | 50.43% | N/A | 0.00% |
| Nano-banana Pro (image) | 14.91% | 28.40% | 46.47% | N/A | 0.00% |
| GPT-4o-image (image) | 69.54% | 59.69% | 19.18% | N/A | 0.00% |
| GPT-image-1.5 (image) | 100.00% | 89.14% | 14.62% | N/A | 0.00% |
| Qwen-image (image) | 28.58% | 7.41% | 73.74% | N/A | 0.00% |
Grid Size: 9×9 — Hard
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 68.80% | 99.60% | 13.45% | 70.40% | 0.00% |
| Sora-2 (video) | 92.86% | 92.86% | 8.65% | 35.71% | 7.14% |
| Wan-2.2 (video) | 93.00% | 100.00% | 16.86% | 3.00% | 0.00% |
| Nano-banana (image) | 27.25% | 38.16% | 41.11% | N/A | 0.00% |
| Nano-banana Pro (image) | 14.75% | 41.36% | 40.11% | N/A | 0.00% |
| GPT-4o-image (image) | 69.80% | 57.66% | 15.76% | N/A | 0.00% |
| GPT-image-1.5 (image) | 100.00% | 87.26% | 12.81% | N/A | 0.00% |
| Qwen-image (image) | 27.26% | 10.44% | 73.59% | N/A | 0.00% |
Table 10: Quantitative results for the Sudoku task using OCR-based evaluation method
Columns: Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑
Grid Size: 4×4 — Easy
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 80.40% | 75.17% | 28.49% | 23.20% | 0.00% |
| Sora-2 (video) | 100.00% | 64.80% | 20.97% | 22.45% | 0.00% |
| Wan-2.2 (video) | 100.00% | 91.67% | 5.68% | 18.00% | 0.00% |
| Nano-banana (image) | 0.00% | 26.93% | 73.85% | N/A | 32.80% |
| Nano-banana Pro (image) | 0.00% | 42.50% | 87.35% | N/A | 14.00% |
| GPT-4o-image (image) | 0.00% | 50.15% | 57.63% | N/A | 7.00% |
| GPT-image-1.5 (image) | 63.33% | 32.06% | 61.90% | N/A | 22.67% |
| Qwen-image (image) | 97.33% | 95.72% | 3.84% | N/A | 0.00% |
Grid Size: 4×4 — Medium
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 86.00% | 90.50% | 25.47% | 35.60% | 0.00% |
| Sora-2 (video) | 100.00% | 59.52% | 25.20% | 33.33% | 0.00% |
| Wan-2.2 (video) | 100.00% | 93.17% | 7.49% | 10.00% | 0.00% |
| Nano-banana (image) | 0.00% | 35.97% | 53.96% | N/A | 13.60% |
| Nano-banana Pro (image) | 0.00% | 43.33% | 85.62% | N/A | 10.00% |
| GPT-4o-image (image) | 0.00% | 56.06% | 41.47% | N/A | 0.75% |
| GPT-image-1.5 (image) | 60.67% | 37.39% | 50.99% | N/A | 10.67% |
| Qwen-image (image) | 92.67% | 99.22% | 2.53% | N/A | 0.00% |
Grid Size: 4×4 — Hard
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 90.80% | 91.53% | 26.02% | 31.20% | 0.00% |
| Sora-2 (video) | 100.00% | 61.41% | 21.96% | 32.61% | 0.00% |
| Wan-2.2 (video) | 100.00% | 94.83% | 11.91% | 4.00% | 0.00% |
| Nano-banana (image) | 0.00% | 49.87% | 42.92% | N/A | 5.20% |
| Nano-banana Pro (image) | 0.00% | 36.00% | 86.86% | N/A | 18.00% |
| GPT-4o-image (image) | 0.00% | 58.29% | 36.80% | N/A | 0.25% |
| GPT-image-1.5 (image) | 81.33% | 39.11% | 39.73% | N/A | 7.33% |
| Qwen-image (image) | 91.33% | 99.28% | 2.92% | N/A | 0.00% |
Grid Size: 9×9 — Easy
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 88.00% | 99.91% | 5.75% | 92.40% | 0.00% |
| Sora-2 (video) | 100.00% | 99.84% | 6.63% | 55.32% | 0.00% |
| Wan-2.2 (video) | 100.00% | 100.00% | 3.72% | 96.00% | 0.00% |
| Nano-banana (image) | 0.00% | 96.07% | 15.70% | N/A | 0.00% |
| Nano-banana Pro (image) | 0.00% | 55.22% | 45.68% | N/A | 0.00% |
| GPT-4o-image (image) | 0.00% | 87.25% | 11.13% | N/A | 0.00% |
| GPT-image-1.5 (image) | 100.00% | 88.79% | 13.18% | N/A | 0.00% |
| Qwen-image (image) | 97.33% | 100.00% | 0.11% | N/A | 0.00% |
Grid Size: 9×9 — Medium
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 84.40% | 99.94% | 6.70% | 94.00% | 0.00% |
| Sora-2 (video) | 100.00% | 100.00% | 6.92% | 44.00% | 0.00% |
| Wan-2.2 (video) | 100.00% | 100.00% | 3.03% | 68.00% | 0.00% |
| Nano-banana (image) | 0.00% | 96.37% | 13.70% | N/A | 0.00% |
| Nano-banana Pro (image) | 0.00% | 53.58% | 35.92% | N/A | 0.00% |
| GPT-4o-image (image) | 0.00% | 85.93% | 10.47% | N/A | 0.00% |
| GPT-image-1.5 (image) | 100.00% | 89.14% | 11.96% | N/A | 0.00% |
| Qwen-image (image) | 92.67% | 100.00% | 0.11% | N/A | 0.00% |
Grid Size: 9×9 — Hard
| Model | Clues Changed ↓ | Constraints Violation ↓ | Completion Accuracy ↑ | Action Reflection ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Veo-3 (video) | 81.60% | 99.99% | 6.93% | 86.40% | 0.00% |
| Sora-2 (video) | 100.00% | 100.00% | 5.32% | 50.00% | 0.00% |
| Wan-2.2 (video) | 100.00% | 100.00% | 3.32% | 50.00% | 0.00% |
| Nano-banana (image) | 0.00% | 95.36% | 13.12% | N/A | 0.00% |
| Nano-banana Pro (image) | 0.00% | 59.88% | 32.55% | N/A | 0.00% |
| GPT-4o-image (image) | 0.00% | 86.25% | 9.79% | N/A | 0.00% |
| GPT-image-1.5 (image) | 100.00% | 87.26% | 12.81% | N/A | 0.00% |
| Qwen-image (image) | 91.33% | 100.00% | 0.08% | N/A | 0.00% |
Note (Figure 11 diagnostic examples, Veo-3): sudoku_easy_4x4_000.mp4 — 91.67% constraint violation, 0% completion accuracy, clues changed at [1,3] from 2→7 in frames 62–64; sudoku_easy_9x9_050.mp4 — 168 clue instances modified, 100% constraint violation, 4.44% completion accuracy, action reflection = 1.
T3 — ARC
Table 11: Distribution of 456 ARC cases across shape consistency and difficulty levels, separated by v1 and v2
Percentages indicate the proportion within each shape consistency group for the corresponding benchmark version.
| Version | Shape Consistency | Easy | Medium | Hard | Total |
|---|---|---|---|---|---|
| V1 | Match | 102 (38.8%) | 124 (47.1%) | 37 (14.1%) | 263 |
| V1 | Mismatch | 34 (28.8%) | 57 (48.3%) | 27 (22.9%) | 118 |
| V1 | Total | 136 | 181 | 64 | 381 |
| V2 | Match | 1 (1.9%) | 27 (50.9%) | 25 (47.2%) | 53 |
| V2 | Mismatch | / | 7 (31.8%) | 15 (68.2%) | 22 |
| V2 | Total | 1 | 34 | 40 | 75 |
| Overall | Match | 103 (31.0%) | 151 (45.4%) | 62 (23.6%) | 316 |
| Overall | Mismatch | 34 (25.0%) | 64 (45.7%) | 42 (29.3%) | 140 |
| Overall | Total | 137 (30.0%) | 215 (47.1%) | 104 (22.8%) | 456 |
Table 12: Final task-level results for the ARC task
Scores match the ARC row in Table 4; subsequent v1/v2 tables provide diagnostic split analyses rather than the final aggregation.
| Model | Type | Final score ↑ |
|---|---|---|
| Veo-3 | Video | 4.80% |
| Sora-2 | Video | 11.67% |
| Wan-2.2 | Video | 0.15% |
| Nano-banana | Image | 8.15% |
| Nano-banana Pro | Image | 28.07% |
| GPT-4o-image | Image | 0.00% |
| GPT-image-1.5 | Image | 12.72% |
| Qwen-image | Image | 2.15% |
Table 13: Diagnostic quantitative results for the ARC v1 split (381 cases)
(GPT-image-1.5 is not listed in Table 13.)
| Model | Pattern Recog. ↑ | Grid Integrity ↑ | Color Accuracy ↑ | Overall ↑ |
|---|---|---|---|---|
| Video Models | ||||
| Veo-3 | 17.32% | 32.98% | 8.22% | 5.16% |
| Sora-2 | 71.99% | 94.58% | 36.75% | 20.18% |
| Wan-2.2 | 0.61% | 13.04% | 0.17% | 0.17% |
| Image Models | ||||
| Nano-banana | 28.42% | 55.79% | 12.63% | 9.21% |
| Nano-banana Pro | 61.98% | 84.73% | 40.42% | 30.54% |
| GPT-4o-image | 1.05% | 10.24% | 0.52% | 0.00% |
| Qwen-image | 1.31% | 4.46% | 0.52% | 0.52% |
Table 14: Diagnostic quantitative results for the ARC v2 split (75 cases)
| Model | Pattern Recog. ↑ | Grid Integrity ↑ | Color Accuracy ↑ | Overall ↑ |
|---|---|---|---|---|
| Video Models | ||||
| Veo-3 | 17.78% | 31.11% | 6.22% | 4.00% |
| Sora-2 | 4.00% | 16.00% | 1.33% | 1.33% |
| Wan-2.2 | 0.00% | 5.78% | 0.00% | 0.00% |
| Image Models | ||||
| Nano-banana | 18.67% | 42.67% | 8.00% | 2.67% |
| Nano-banana Pro | 62.50% | 83.93% | 44.64% | 30.36% |
| GPT-4o-image | 1.33% | 2.67% | 1.33% | 0.00% |
| Qwen-image | 1.33% | 5.33% | 1.33% | 1.33% |
Table 15: Diagnostic quantitative breakdown results (Match and Mismatch) for the ARC v1 task (381 cases)
| Model / Category | Pattern Recog. ↑ | Grid Integrity ↑ | Color Accuracy ↑ | Overall ↑ |
|---|---|---|---|---|
| Video Models | ||||
| Veo-3 — Match | 17.74% | 35.74% | 9.51% | 5.70% |
| Veo-3 — Mismatch | 16.38% | 26.84% | 5.37% | 3.95% |
| Sora-2 — Match | 67.29% | 93.46% | 35.05% | 17.76% |
| Sora-2 — Mismatch | 80.51% | 96.61% | 39.83% | 24.58% |
| Wan-2.2 — Match | 0.51% | 13.31% | 0.00% | 0.00% |
| Wan-2.2 — Mismatch | 0.85% | 12.43% | 0.56% | 0.56% |
| Image Models | ||||
| Nano-banana — Match | 24.05% | 53.82% | 11.07% | 8.40% |
| Nano-banana — Mismatch | 38.14% | 60.17% | 16.10% | 11.02% |
| Nano-banana Pro — Match | 62.24% | 89.63% | 42.74% | 31.54% |
| Nano-banana Pro — Mismatch | 61.29% | 72.04% | 34.41% | 27.96% |
| GPT-4o-image — Match | 0.38% | 9.51% | 0.76% | 0.00% |
| GPT-4o-image — Mismatch | 2.54% | 11.86% | 0.00% | 0.00% |
| Qwen-image — Match | 0.76% | 4.56% | 0.76% | 0.76% |
| Qwen-image — Mismatch | 2.54% | 4.24% | 0.00% | 0.00% |
Table 16: Diagnostic quantitative breakdown results for the ARC v1 task (381 cases) across difficulty levels (Easy, Medium, Hard)
| Model / Difficulty | Pattern Recog. ↑ | Grid Integrity ↑ | Color Accuracy ↑ | Overall ↑ |
|---|---|---|---|---|
| Veo-3: Match | ||||
| Easy | 18.95% | 40.52% | 11.76% | 5.88% |
| Medium | 17.74% | 35.48% | 8.06% | 5.65% |
| Hard | 14.41% | 23.42% | 8.11% | 5.41% |
| Veo-3: Mismatch | ||||
| Easy | 19.61% | 37.25% | 6.86% | 4.90% |
| Medium | 14.04% | 23.39% | 5.26% | 3.51% |
| Hard | 17.28% | 20.99% | 3.70% | 3.70% |
| Sora-2: Match | ||||
| Easy | 73.17% | 90.24% | 42.68% | 19.51% |
| Medium | 63.81% | 95.24% | 33.33% | 19.05% |
| Hard | 62.96% | 96.30% | 18.52% | 7.41% |
| Sora-2: Mismatch | ||||
| Easy | 85.29% | 91.18% | 41.18% | 29.41% |
| Medium | 80.70% | 100.00% | 36.84% | 19.30% |
| Hard | 74.07% | 96.30% | 44.44% | 29.63% |
| Wan-2.2: Match | ||||
| Easy | 0.00% | 16.34% | 0.00% | 0.00% |
| Medium | 0.54% | 12.37% | 0.00% | 0.00% |
| Hard | 1.80% | 8.11% | 0.00% | 0.00% |
| Wan-2.2: Mismatch | ||||
| Easy | 0.00% | 17.65% | 0.00% | 0.00% |
| Medium | 1.75% | 9.36% | 1.17% | 1.17% |
| Hard | 0.00% | 12.35% | 0.00% | 0.00% |
| Nano-banana: Match | ||||
| Easy | 27.45% | 56.86% | 13.73% | 8.82% |
| Medium | 20.33% | 50.41% | 8.13% | 7.32% |
| Hard | 27.03% | 56.76% | 13.51% | 10.81% |
| Nano-banana: Mismatch | ||||
| Easy | 47.06% | 85.29% | 11.76% | 11.76% |
| Medium | 35.09% | 52.63% | 19.30% | 10.53% |
| Hard | 33.33% | 44.44% | 14.81% | 11.11% |
| Nano-banana Pro: Match | ||||
| Easy | 62.77% | 90.43% | 45.74% | 31.91% |
| Medium | 60.71% | 89.29% | 37.50% | 29.46% |
| Hard | 65.71% | 88.57% | 51.43% | 37.14% |
| Nano-banana Pro: Mismatch | ||||
| Easy | 60.71% | 71.43% | 35.71% | 25.00% |
| Medium | 65.12% | 74.42% | 37.21% | 30.23% |
| Hard | 54.55% | 68.18% | 27.27% | 27.27% |
| GPT-4o-image: Match | ||||
| Easy | 0.98% | 10.78% | 1.96% | 0.00% |
| Medium | 0.00% | 11.29% | 0.00% | 0.00% |
| Hard | 0.00% | 0.00% | 0.00% | 0.00% |
| GPT-4o-image: Mismatch | ||||
| Easy | 0.00% | 11.76% | 0.00% | 0.00% |
| Medium | 1.75% | 10.53% | 0.00% | 0.00% |
| Hard | 7.41% | 14.81% | 0.00% | 0.00% |
| Qwen-image: Match | ||||
| Easy | 0.98% | 5.88% | 0.98% | 0.98% |
| Medium | 0.00% | 3.23% | 0.00% | 0.00% |
| Hard | 2.70% | 5.41% | 2.70% | 2.70% |
| Qwen-image: Mismatch | ||||
| Easy | 2.94% | 2.94% | 0.00% | 0.00% |
| Medium | 3.51% | 7.02% | 0.00% | 0.00% |
| Hard | 0.00% | 0.00% | 0.00% | 0.00% |
Table 17: Diagnostic quantitative breakdown results (Match and Mismatch) for the ARC v2 task (75 cases)
| Model / Category | Pattern Recog. ↑ | Grid Integrity ↑ | Color Accuracy ↑ | Overall ↑ |
|---|---|---|---|---|
| Video Models | ||||
| Veo-3 — Match | 17.61% | 32.08% | 6.92% | 3.77% |
| Veo-3 — Mismatch | 18.18% | 28.79% | 4.55% | 4.55% |
| Sora-2 — Match | 5.66% | 18.87% | 1.89% | 1.89% |
| Sora-2 — Mismatch | 0.00% | 9.09% | 0.00% | 0.00% |
| Wan-2.2 — Match | 0.00% | 4.40% | 0.00% | 0.00% |
| Wan-2.2 — Mismatch | 0.00% | 9.09% | 0.00% | 0.00% |
| Image Models | ||||
| Nano-banana — Match | 18.87% | 45.28% | 11.32% | 3.77% |
| Nano-banana — Mismatch | 18.18% | 36.36% | 0.00% | 0.00% |
| Nano-banana Pro — Match | 64.29% | 88.10% | 50.00% | 33.33% |
| Nano-banana Pro — Mismatch | 57.14% | 71.43% | 28.57% | 21.43% |
| GPT-4o-image — Match | 0.00% | 1.89% | 0.00% | 0.00% |
| GPT-4o-image — Mismatch | 4.55% | 4.55% | 4.55% | 0.00% |
| Qwen-image — Match | 1.89% | 5.66% | 1.89% | 1.89% |
| Qwen-image — Mismatch | 0.00% | 4.55% | 0.00% | 0.00% |
Table 18: Diagnostic quantitative breakdown results for the ARC v2 task (75 cases) across difficulty levels (Easy, Medium, Hard)
* The Easy level of ARC v2 contains only one evaluation case; therefore, a model solving this case correctly achieves a 100% score.
| Model / Difficulty | Pattern Recog. ↑ | Grid Integrity ↑ | Color Accuracy ↑ | Overall ↑ |
|---|---|---|---|---|
| Veo-3: Match | ||||
| Easy | 0.00% | 33.33% | 0.00% | 0.00% |
| Medium | 20.99% | 34.57% | 8.64% | 6.17% |
| Hard | 14.67% | 29.33% | 5.33% | 1.33% |
| Veo-3: Mismatch | ||||
| Medium | 23.81% | 42.86% | 4.76% | 4.76% |
| Hard | 15.56% | 22.22% | 4.44% | 4.44% |
| Sora-2: Match | ||||
| Easy | 0.00% | 0.00% | 0.00% | 0.00% |
| Medium | 7.41% | 22.22% | 3.70% | 3.70% |
| Hard | 4.00% | 16.00% | 0.00% | 0.00% |
| Sora-2: Mismatch | ||||
| Medium | 0.00% | 28.57% | 0.00% | 0.00% |
| Hard | 0.00% | 0.00% | 0.00% | 0.00% |
| Wan-2.2: Match | ||||
| Easy | 0.00% | 33.33% | 0.00% | 0.00% |
| Medium | 0.00% | 4.94% | 0.00% | 0.00% |
| Hard | 0.00% | 2.67% | 0.00% | 0.00% |
| Wan-2.2: Mismatch | ||||
| Medium | 0.00% | 9.52% | 0.00% | 0.00% |
| Hard | 0.00% | 8.89% | 0.00% | 0.00% |
| Nano-banana: Match | ||||
| Easy | 100.00% | 100.00% | 100.00% | 100.00%* |
| Medium | 18.52% | 55.56% | 7.41% | 0.00% |
| Hard | 16.00% | 32.00% | 12.00% | 4.00% |
| Nano-banana: Mismatch | ||||
| Medium | 42.86% | 57.14% | 0.00% | 0.00% |
| Hard | 6.67% | 26.67% | 0.00% | 0.00% |
| Nano-banana Pro: Match | ||||
| Easy | 100.00% | 100.00% | 100.00% | 100.00%* |
| Medium | 66.67% | 95.24% | 61.90% | 38.10% |
| Hard | 60.00% | 80.00% | 35.00% | 25.00% |
| Nano-banana Pro: Mismatch | ||||
| Medium | 80.00% | 100.00% | 40.00% | 20.00% |
| Hard | 44.44% | 55.56% | 22.22% | 22.22% |
| GPT-4o-image: Match | ||||
| Easy | 0.00% | 0.00% | 0.00% | 0.00% |
| Medium | 0.00% | 3.70% | 0.00% | 0.00% |
| Hard | 0.00% | 0.00% | 0.00% | 0.00% |
| GPT-4o-image: Mismatch | ||||
| Medium | 0.00% | 0.00% | 0.00% | 0.00% |
| Hard | 6.67% | 6.67% | 6.67% | 0.00% |
| Qwen-image: Match | ||||
| Easy | 0.00% | 0.00% | 0.00% | 0.00% |
| Medium | 3.70% | 7.41% | 3.70% | 3.70% |
| Hard | 0.00% | 4.00% | 0.00% | 0.00% |
| Qwen-image: Mismatch | ||||
| Medium | 0.00% | 0.00% | 0.00% | 0.00% |
| Hard | 0.00% | 6.67% | 0.00% | 0.00% |
Additional narrative numbers (E.4)
- Final task-level hierarchy: Nano-banana Pro 28.07%, GPT-image-1.5 12.72%, Sora-2 (strongest video) 11.67%.
- Nano-banana Pro Overall — Easy 30.89%, Medium 30.39%, Hard 30.23%; Grid Integrity 86.18% → 86.74% → 77.91%. Sora-2 Overall — Easy 22.22%, Medium 16.33%, Hard 10.64%; Pattern Recognition 76.07% → 40.43%. Veo-3 ~5% across levels; Grid Integrity 39.66% → 32.40% → 24.04%.
- v1→v2 collapse: Sora-2 20.18% → 1.33% (93% decline); Nano-banana 9.21% → 2.67% (71% decline); Nano-banana Pro 30.54% → 30.36% (stable). Veo-3 Grid Integrity 32.98% (v1) vs 31.11% (v2).
- Metric cascade (Nano-banana v1): Grid Integrity 55.79% → Pattern Recognition 28.42% → Color Accuracy 12.63% → Overall 9.21%. Color Accuracy = 20–25% of Grid Integrity across all models.
T4 — Math
Table 19: Per-benchmark sample counts
| Dataset | Level | Sample Count |
|---|---|---|
| GSM8K | Grade School | 50 |
| MATH500 | High School Competition | 50 |
| AIME 2024 | Invitational Competition | 30 |
| AIME 2025 | Invitational Competition | 30 |
| Omni-MATH | Multi-level (T0-T4) | 167 |
| Total | – | 327 |
Table 20: Omni-MATH sample distribution across difficulty levels × categories
| Difficulty | Algebra | Applied Math | Calculus | Discrete Math | Geometry | Precalculus | Number | Other |
|---|---|---|---|---|---|---|---|---|
| T0 (Easiest) | 4 | 5 | 5 | 5 | 5 | 5 | 4 | 1 |
| T1 | 5 | 5 | 4 | 5 | 4 | 5 | 5 | 0 |
| T2 | 5 | 5 | 5 | 5 | 5 | 5 | 5 | 1 |
| T3 | 5 | 5 | 5 | 5 | 5 | 4 | 5 | 0 |
| T4 (Hardest) | 5 | 5 | 1 | 5 | 5 | 4 | 5 | 0 |
| Total | 24 | 25 | 20 | 25 | 24 | 23 | 24 | 2 |
(Omni-MATH category totals sum to 167.)
Table 21: Final task-level results for the Math task
| Model | Type | Final score ↑ |
|---|---|---|
| Veo-3 | Video | 10.61% |
| Sora-2 | Video | 10.56% |
| Wan-2.2 | Video | 0.00% |
| Nano-banana | Image | 11.47% |
| Nano-banana Pro | Image | 72.69% |
| GPT-4o-image | Image | 26.24% |
| GPT-image-1.5 | Image | 25.52% |
| Qwen-image | Image | 11.22% |
| Gemini-3-Flash | Text baseline | 71.38% |
| Gemini-3-Pro | Text baseline | 74.14% |
Table 22: Diagnostic fine-grained results for the Math task across the original benchmark splits
Columns: Process Success Rate ↑ | Outcome Success Rate ↑ | Action Reflection ↑ | Overall Success Rate ↑ (Primary Metric). "–" = missing per-dataset result.
Dataset: GSM8K
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 12.00% | 74.00% | 12.00% | 12.00% |
| Sora-2 (Video) | 38.00% | 64.00% | 16.00% | 30.00% |
| Wan-2.2 (Video) | 2.00% | 2.00% | 0.00% | 2.00% |
| Nano-banana (Image) | 44.00% | 88.00% | N/A | 42.00% |
| Nano-banana Pro (Image) | 97.83% | 97.83% | N/A | 97.83% |
| GPT-4o-image (Image) | 80.00% | 83.48% | N/A | 75.65% |
| Qwen-image (Image) | 44.00% | 44.00% | N/A | 44.00% |
Dataset: MATH500
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 20.00% | 52.00% | 14.00% | 18.00% |
| Sora-2 (Video) | 34.04% | 59.57% | 10.64% | 31.91% |
| Wan-2.2 (Video) | 3.33% | 6.00% | 0.00% | 3.33% |
| Nano-banana (Image) | 16.00% | 74.00% | N/A | 16.00% |
| Nano-banana Pro (Image) | 91.84% | 91.84% | N/A | 91.84% |
| GPT-4o-image (Image) | 29.14% | 42.45% | N/A | 27.34% |
| Qwen-image (Image) | – | – | N/A | – |
Dataset: AIME24
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 8.33% | 8.33% | 5.00% | 1.67% |
| Sora-2 (Video) | 8.70% | 13.04% | 4.35% | 4.35% |
| Wan-2.2 (Video) | 15.56% | 22.22% | 11.11% | 15.56% |
| Nano-banana (Image) | 5.00% | 15.00% | N/A | 0.00% |
| Nano-banana Pro (Image) | 63.64% | 36.36% | N/A | 31.82% |
| GPT-4o-image (Image) | 3.33% | 10.00% | N/A | 0.00% |
| Qwen-image (Image) | 1.12% | 1.12% | N/A | 1.12% |
Dataset: AIME25
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 3.33% | 11.67% | 1.67% | 3.33% |
| Sora-2 (Video) | 8.70% | 21.74% | 4.35% | 0.00% |
| Wan-2.2 (Video) | 5.56% | 10.00% | 3.33% | 5.56% |
| Nano-banana (Image) | 1.75% | 33.33% | N/A | 1.75% |
| Nano-banana Pro (Image) | 71.43% | 90.48% | N/A | 66.67% |
| GPT-4o-image (Image) | 0.00% | 3.33% | N/A | 0.00% |
| Qwen-image (Image) | 4.00% | 4.00% | N/A | 4.00% |
Dataset: Omni-MATH
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 4.79% | 15.57% | 5.09% | 3.89% |
| Sora-2 (Video) | 0.62% | 1.88% | 6.88% | 0.62% |
| Wan-2.2 (Video) | 0.41% | 3.46% | 0.61% | 0.41% |
| Nano-banana (Image) | 3.90% | 39.94% | N/A | 3.90% |
| Nano-banana Pro (Image) | 65.77% | 85.59% | N/A | 63.06% |
| GPT-4o-image (Image) | 4.79% | 41.92% | N/A | 4.79% |
| Qwen-image (Image) | 6.67% | 6.67% | N/A | 6.67% |
Table 23: Quantitative breakdown results for the Omni-MATH task (five difficulty levels T0–T4)
Columns: Process Success Rate ↑ | Outcome Success Rate ↑ | Action Reflection ↑ | Overall Success Rate ↑ (Primary Metric).
T0
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 6.06% | 12.12% | 3.03% | 6.06% |
| Sora-2 (Video) | 3.03% | 6.06% | 15.15% | 3.03% |
| Wan-2.2 (Video) | 0.98% | 4.90% | 2.94% | 0.98% |
| Nano-banana (Image) | 0.00% | 33.82% | N/A | 0.00% |
| Nano-banana Pro (Image) | 53.85% | 84.62% | N/A | 53.85% |
| GPT-4o-image (Image) | 0.00% | 0.00% | N/A | 0.00% |
| Qwen-image (Image) | 22.00% | 22.00% | N/A | 22.00% |
T1
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 4.55% | 4.55% | 1.52% | 3.03% |
| Sora-2 (Video) | 0.00% | 3.12% | 3.12% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 6.06% | 0.00% | 0.00% |
| Nano-banana (Image) | 0.00% | 19.70% | N/A | 0.00% |
| Nano-banana Pro (Image) | 26.67% | 60.00% | N/A | 26.67% |
| GPT-4o-image (Image) | 0.00% | 0.00% | N/A | 0.00% |
| Qwen-image (Image) | 8.42% | 13.68% | N/A | 8.42% |
T2
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 5.56% | 16.67% | 2.78% | 5.56% |
| Sora-2 (Video) | 0.00% | 0.00% | 2.86% | 0.00% |
| Wan-2.2 (Video) | 0.93% | 3.70% | 0.00% | 0.93% |
| Nano-banana (Image) | 4.17% | 47.22% | N/A | 4.17% |
| Nano-banana Pro (Image) | 70.00% | 96.67% | N/A | 66.67% |
| GPT-4o-image (Image) | 2.78% | 2.78% | N/A | 2.78% |
| Qwen-image (Image) | 7.92% | 8.91% | N/A | 7.92% |
T3
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 1.96% | 13.24% | 1.47% | 1.47% |
| Sora-2 (Video) | 0.00% | 0.00% | 9.68% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 1.96% | 0.00% | 0.00% |
| Nano-banana (Image) | 5.88% | 45.59% | N/A | 5.88% |
| Nano-banana Pro (Image) | 70.83% | 79.17% | N/A | 62.50% |
| GPT-4o-image (Image) | 0.00% | 2.86% | N/A | 0.00% |
| Qwen-image (Image) | 16.67% | 12.22% | N/A | 12.22% |
T4
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 5.17% | 36.67% | 15.00% | 5.00% |
| Sora-2 (Video) | 0.00% | 0.00% | 3.45% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 0.00% | 0.00% | 0.00% |
| Nano-banana (Image) | 9.52% | 53.97% | N/A | 9.52% |
| Nano-banana Pro (Image) | 82.76% | 93.10% | N/A | 82.76% |
| GPT-4o-image (Image) | 0.00% | 6.67% | N/A | 0.00% |
| Qwen-image (Image) | 25.68% | 22.97% | N/A | 22.97% |
Table 24: Quantitative breakdown results for the Omni-MATH task (eight categories)
Columns: Process Success Rate ↑ | Outcome Success Rate ↑ | Action Reflection ↑ | Overall Success Rate ↑ (Primary Metric).
Algebra
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 4.17% | 12.50% | 4.17% | 4.17% |
| Sora-2 (Video) | 4.35% | 8.70% | 13.04% | 4.35% |
| Wan-2.2 (Video) | 0.00% | 5.56% | 0.00% | 0.00% |
| Nano-banana (Image) | 4.17% | 45.83% | N/A | 4.17% |
| Nano-banana Pro (Image) | 61.11% | 83.33% | N/A | 55.56% |
| GPT-4o-image (Image) | 0.00% | 4.00% | N/A | 0.00% |
| Qwen-image (Image) | 19.70% | 24.24% | N/A | 19.70% |
Applied Math
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 0.00% | 20.00% | 8.00% | 0.00% |
| Sora-2 (Video) | 0.00% | 0.00% | 8.00% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 5.33% | 0.00% | 0.00% |
| Nano-banana (Image) | 16.00% | 32.00% | N/A | 16.00% |
| Nano-banana Pro (Image) | 68.75% | 87.50% | N/A | 68.75% |
| GPT-4o-image (Image) | 4.00% | 4.00% | N/A | 4.00% |
| Qwen-image (Image) | 14.49% | 14.49% | N/A | 14.49% |
Calculus
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 5.00% | 10.00% | 5.00% | 0.00% |
| Sora-2 (Video) | 0.00% | 0.00% | 10.00% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 1.67% | 0.00% | 0.00% |
| Nano-banana (Image) | 0.00% | 35.00% | N/A | 0.00% |
| Nano-banana Pro (Image) | 66.67% | 80.00% | N/A | 53.33% |
| GPT-4o-image (Image) | 0.00% | 0.00% | N/A | 0.00% |
| Qwen-image (Image) | 15.79% | 15.79% | N/A | 15.79% |
Discrete Math
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 0.00% | 12.00% | 4.00% | 0.00% |
| Sora-2 (Video) | 0.00% | 0.00% | 4.17% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 5.33% | 0.00% | 0.00% |
| Nano-banana (Image) | 0.00% | 24.00% | N/A | 0.00% |
| Nano-banana Pro (Image) | 54.55% | 72.73% | N/A | 54.55% |
| GPT-4o-image (Image) | 0.00% | 0.00% | N/A | 0.00% |
| Qwen-image (Image) | 15.94% | 17.39% | N/A | 15.94% |
Geometry
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 8.33% | 25.00% | 4.17% | 8.33% |
| Sora-2 (Video) | 0.00% | 0.00% | 0.00% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 1.39% | 4.17% | 0.00% |
| Nano-banana (Image) | 0.00% | 45.83% | N/A | 0.00% |
| Nano-banana Pro (Image) | 78.57% | 92.86% | N/A | 78.57% |
| GPT-4o-image (Image) | 0.00% | 0.00% | N/A | 0.00% |
| Qwen-image (Image) | 23.64% | 20.00% | N/A | 18.18% |
Precalculus
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 4.35% | 21.74% | 8.70% | 0.00% |
| Sora-2 (Video) | 0.00% | 0.00% | 0.00% | 0.00% |
| Wan-2.2 (Video) | 0.00% | 1.45% | 0.00% | 0.00% |
| Nano-banana (Image) | 13.04% | 65.22% | N/A | 13.04% |
| Nano-banana Pro (Image) | 68.18% | 86.36% | N/A | 68.18% |
| GPT-4o-image (Image) | 0.00% | 4.17% | N/A | 0.00% |
| Qwen-image (Image) | 20.90% | 16.42% | N/A | 16.42% |
Number
| Model | Process SR | Outcome SR | Action Reflection | Overall SR |
|---|---|---|---|---|
| Veo-3 (Video) | 4.17% | 8.33% | 12.50% | 4.17% |
| Sora-2 (Video) | 0.00% | 0.00% | 0.00% | 0.00% |
| Wan-2.2 (Video) | 1.39% | 1.39% | 0.00% | 1.39% |
| Nano-banana (Image) | 0.00% | 50.00% | N/A | 0.00% |
| Nano-banana Pro (Image) | 64.29% | 92.86% | N/A | 64.29% |
| GPT-4o-image (Image) | 0.00% | 4.00% | N/A | 0.00% |
| Qwen-image (Image) | 2.82% | 4.23% | N/A | 2.82% |
(Note: Table 24 in the source ends at the "Number" category; no "Other" category rows appear in the transcribed table.)
T5–T8 — Embodied Navigation (shared tables)
Table 25: Distribution of evaluation samples across the 24 hard-level configurations
Every one of the 24 configurations holds 5 samples for each of the four navigation tasks; each task totals 120.
| Env. Complexity | View Fidelity | Distance | Destination Type | Last-Mile Nav. | Top-down View Nav. | 3D R.-W. Nav. | SLAG |
|---|---|---|---|---|---|---|---|
| 1 Floor | quality03 | short | color mark | 5 | 5 | 5 | 5 |
| 1 Floor | quality03 | short | location description | 5 | 5 | 5 | 5 |
| 1 Floor | quality03 | long | color mark | 5 | 5 | 5 | 5 |
| 1 Floor | quality03 | long | location description | 5 | 5 | 5 | 5 |
| 1 Floor | quality04 | short | color mark | 5 | 5 | 5 | 5 |
| 1 Floor | quality04 | short | location description | 5 | 5 | 5 | 5 |
| 1 Floor | quality04 | long | color mark | 5 | 5 | 5 | 5 |
| 1 Floor | quality04 | long | location description | 5 | 5 | 5 | 5 |
| 1 Floor | quality05 | short | color mark | 5 | 5 | 5 | 5 |
| 1 Floor | quality05 | short | location description | 5 | 5 | 5 | 5 |
| 1 Floor | quality05 | long | color mark | 5 | 5 | 5 | 5 |
| 1 Floor | quality05 | long | location description | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality03 | short | color mark | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality03 | short | location description | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality03 | long | color mark | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality03 | long | location description | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality04 | short | color mark | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality04 | short | location description | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality04 | long | color mark | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality04 | long | location description | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality05 | short | color mark | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality05 | short | location description | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality05 | long | color mark | 5 | 5 | 5 | 5 |
| 2 Plus Floors | quality05 | long | location description | 5 | 5 | 5 | 5 |
| Total (all 24 configs) | 120 | 120 | 120 | 120 |
Table 26: Final task-level scores for Embodied Navigation (all values %)
| Task | Veo-3 | Sora-2 | Wan-2.2 | Nano-banana | Nano-banana Pro | GPT-4o-image | GPT-image-1.5 | Qwen-image |
|---|---|---|---|---|---|---|---|---|
| Last-Mile Nav. | 60.00% | 0.00% | 14.17% | 74.17% | 75.83% | 0.00% | 55.84% | 16.67% |
| Top-down View Nav. | 19.49% | 3.39% | 5.09% | 11.11% | 33.05% | 3.39% | 26.27% | 5.08% |
| 3D R.-W. Nav. | 22.50% | 0.00% | 24.17% | 79.17% | 85.00% | 13.33% | 77.92% | 38.33% |
| SLAG | 11.02% | 12.50% | 0.85% | 28.79% | 37.29% | 16.67% | 31.36% | 6.78% |
T5 — Panoramic View Last-Mile Navigation
Table 27: Quantitative results for the Panoramic View Last-Mile Navigation benchmark (VLM-based)
Columns: S.S.(3D) | O.S.(3D) | Obj. Sem. | Agent Con. | Spa. Ali. | Des. Inte. | Scene Con. | Succ(3D) Orig. Dest. | Physics Validness | Overall Success. Models: Veo-3, Sora-2 (video); Nano-banana, GPT-4o-image (image).
| Group / Model | S.S.(3D) | O.S.(3D) | Obj. Sem. | Agent Con. | Spa. Ali. | Des. Inte. | Scene Con. | Succ(3D) Orig. Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Env. Complexity — floor01 | ||||||||||
| Veo-3 | 90.00% | 90.00% | 93.33% | 93.33% | 93.33% | 90.00% | 98.33% | 90.00% | 81.67% | 73.33% |
| Sora-2 | 0.00% | 1.67% | 93.33% | 88.33% | 93.33% | 0.00% | 0.00% | 0.00% | 81.67% | 0.00% |
| Nano-banana | 76.67% | 80.00% | 96.67% | 88.33% | 88.33% | 80.00% | 91.67% | 75.00% | 88.33% | 73.33% |
| GPT-4o-image | 0.00% | 0.00% | 100.00% | 1.67% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| Env. Complexity — floor02plus | ||||||||||
| Veo-3 | 58.33% | 66.67% | 95.00% | 93.33% | 91.67% | 55.00% | 91.67% | 55.00% | 83.33% | 46.67% |
| Sora-2 | 0.00% | 0.00% | 81.36% | 77.97% | 76.27% | 0.00% | 1.69% | 0.00% | 64.41% | 0.00% |
| Nano-banana | 80.00% | 80.00% | 96.67% | 90.00% | 86.67% | 80.00% | 90.00% | 78.33% | 86.67% | 75.00% |
| GPT-4o-image | 0.00% | 0.00% | 91.67% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| View Fidelity — quality03 | ||||||||||
| Veo-3 | 72.50% | 80.00% | 92.50% | 92.50% | 90.00% | 70.00% | 92.50% | 70.00% | 77.50% | 55.00% |
| Sora-2 | 0.00% | 2.56% | 79.49% | 82.05% | 76.92% | 0.00% | 2.56% | 0.00% | 69.23% | 0.00% |
| Nano-banana | 62.50% | 62.50% | 92.50% | 87.50% | 82.50% | 62.50% | 80.00% | 62.50% | 82.50% | 60.00% |
| GPT-4o-image | 0.00% | 0.00% | 90.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| View Fidelity — quality04 | ||||||||||
| Veo-3 | 75.00% | 75.00% | 95.00% | 95.00% | 90.00% | 75.00% | 97.50% | 75.00% | 82.50% | 62.50% |
| Sora-2 | 0.00% | 0.00% | 85.00% | 80.00% | 87.50% | 0.00% | 0.00% | 0.00% | 67.50% | 0.00% |
| Nano-banana | 85.00% | 90.00% | 100.00% | 87.50% | 87.50% | 90.00% | 92.50% | 80.00% | 87.50% | 80.00% |
| GPT-4o-image | 0.00% | 0.00% | 100.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| View Fidelity — quality05 | ||||||||||
| Veo-3 | 75.00% | 80.00% | 95.00% | 92.50% | 97.50% | 72.50% | 95.00% | 72.50% | 87.50% | 62.50% |
| Sora-2 | 0.00% | 0.00% | 97.50% | 87.50% | 90.00% | 0.00% | 0.00% | 0.00% | 82.50% | 0.00% |
| Nano-banana | 87.50% | 87.50% | 97.50% | 92.50% | 92.50% | 87.50% | 100.00% | 87.50% | 92.50% | 82.50% |
| GPT-4o-image | 0.00% | 0.00% | 97.50% | 2.50% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| Trajectory Distance — short | ||||||||||
| Veo-3 | 81.67% | 83.33% | 96.67% | 93.33% | 90.00% | 80.00% | 98.33% | 80.00% | 83.33% | 66.67% |
| Sora-2 | 0.00% | 0.00% | 88.33% | 83.33% | 85.00% | 0.00% | 1.67% | 0.00% | 76.67% | 0.00% |
| Nano-banana | 86.67% | 86.67% | 98.33% | 93.33% | 91.67% | 86.67% | 95.00% | 85.00% | 91.67% | 81.67% |
| GPT-4o-image | 0.00% | 0.00% | 95.00% | 1.67% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| Trajectory Distance — long | ||||||||||
| Veo-3 | 66.67% | 73.33% | 91.67% | 93.33% | 95.00% | 65.00% | 91.67% | 65.00% | 81.67% | 53.33% |
| Sora-2 | 0.00% | 1.69% | 86.44% | 83.05% | 84.75% | 0.00% | 0.00% | 0.00% | 69.49% | 0.00% |
| Nano-banana | 70.00% | 73.33% | 95.00% | 85.00% | 83.33% | 73.33% | 86.67% | 68.33% | 83.33% | 66.67% |
| GPT-4o-image | 0.00% | 0.00% | 96.67% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| Destination Spec. — color mark | ||||||||||
| Veo-3 | 83.33% | 88.33% | 96.67% | 93.33% | 93.33% | 80.00% | 93.33% | 80.00% | 85.00% | 70.00% |
| Sora-2 | 0.00% | 1.69% | 86.44% | 77.97% | 81.36% | 0.00% | 0.00% | 0.00% | 67.80% | 0.00% |
| Nano-banana | 76.67% | 80.00% | 95.00% | 85.00% | 85.00% | 81.67% | 80.00% | 86.67% | 73.33% | 81.67% |
| GPT-4o-image | 0.00% | 0.00% | 93.33% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| Destination Spec. — location description | ||||||||||
| Veo-3 | 65.00% | 68.33% | 91.67% | 93.33% | 91.67% | 65.00% | 96.67% | 65.00% | 80.00% | 50.00% |
| Sora-2 | 0.00% | 0.00% | 88.33% | 88.33% | 88.33% | 0.00% | 1.67% | 0.00% | 78.33% | 0.00% |
| Nano-banana | 80.00% | 80.00% | 98.33% | 93.33% | 93.33% | 80.00% | 95.00% | 80.00% | 93.33% | 78.33% |
| GPT-4o-image | 0.00% | 0.00% | 98.33% | 1.67% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
Source note (color-mark Nano-banana row): the raw text lists eleven values (76.67, 80.00, 95.00, 85.00, 85.00, 81.67, 80.00, 86.67, 73.33, 81.67, 70.00); the mapping above follows the ten-column header order and the trailing 70.00 is an ambiguous plain-text extraction artifact.
Table 28: Quantitative Geo-Align results for the Panoramic View Last-Mile Navigation task
Column groups: Task Success (SR, SPL); Path Fidelity (nDTW, sDTW); Trajectory Error (ATE, RPE); Path Length (Pred, GT); Navigation Error (Nav, Dir). Models: Veo-3, Sora-2, Wan-2.2. (Table 28 omits a quality04 view-fidelity level.)
| Group / Model | SR | SPL | nDTW | sDTW | ATE | RPE | Pred | GT | Nav | Dir |
|---|---|---|---|---|---|---|---|---|---|---|
| Env. Complexity — floor01 | ||||||||||
| Veo-3 | 83.3% | 50.6% | 53.9% | 48.7% | 0.91 | 0.37 | 6.80 | 3.56 | 0.66 | 11.18 |
| Sora-2 | 55.0% | 26.5% | 41.7% | 29.0% | 1.70 | 0.60 | 11.43 | 3.56 | 1.40 | 14.02 |
| Wan-2.2 | 53.3% | 31.7% | 44.5% | 32.6% | 2.65 | 0.77 | 16.48 | 3.56 | 1.46 | 24.06 |
| Env. Complexity — floor02plus | ||||||||||
| Veo-3 | 50.0% | 32.3% | 41.1% | 28.7% | 1.46 | 0.48 | 8.61 | 3.75 | 1.60 | 19.47 |
| Sora-2 | 25.0% | 9.9% | 32.0% | 13.7% | 1.85 | 0.59 | 10.90 | 3.75 | 2.00 | 17.13 |
| Wan-2.2 | 33.9% | 14.9% | 34.5% | 17.7% | 1.74 | 0.60 | 11.00 | 3.75 | 1.61 | 16.61 |
| View Fidelity — quality03 | ||||||||||
| Veo-3 | 62.5% | 42.5% | 46.0% | 35.3% | 1.20 | 0.37 | 7.98 | 4.43 | 1.10 | 12.39 |
| Sora-2 | 22.5% | 10.7% | 32.2% | 11.8% | 2.09 | 0.57 | 12.60 | 4.43 | 2.14 | 16.71 |
| Wan-2.2 | 37.5% | 22.5% | 38.3% | 21.5% | 3.41 | 0.78 | 19.82 | 4.43 | 1.68 | 17.28 |
| View Fidelity — quality05 | ||||||||||
| Veo-3 | 77.5% | 46.8% | 53.2% | 47.2% | 0.93 | 0.39 | 6.53 | 3.08 | 0.83 | 14.28 |
| Sora-2 | 50.0% | 23.0% | 40.8% | 28.1% | 1.64 | 0.63 | 10.98 | 3.08 | 1.59 | 14.80 |
| Wan-2.2 | 52.5% | 29.2% | 45.0% | 31.8% | 1.41 | 0.53 | 9.01 | 3.08 | 1.27 | 19.79 |
| Trajectory Complexity — noturn | ||||||||||
| Veo-3 | 67.2% | 37.5% | 49.0% | 39.4% | 1.06 | 0.46 | 6.83 | 2.84 | 0.91 | 14.98 |
| Sora-2 | 48.3% | 21.4% | 40.8% | 26.9% | 1.50 | 0.57 | 9.06 | 2.84 | 1.48 | 16.13 |
| Wan-2.2 | 53.4% | 28.2% | 43.9% | 31.4% | 1.31 | 0.57 | 8.54 | 2.84 | 1.25 | 18.55 |
| Trajectory Complexity — oneturn | ||||||||||
| Veo-3 | 67.2% | 46.1% | 46.4% | 38.8% | 1.29 | 0.38 | 8.52 | 4.47 | 1.32 | 15.38 |
| Sora-2 | 32.8% | 15.6% | 33.3% | 16.4% | 2.04 | 0.61 | 13.28 | 4.47 | 1.91 | 14.92 |
| Wan-2.2 | 34.5% | 18.9% | 35.5% | 19.4% | 3.11 | 0.81 | 19.13 | 4.47 | 1.81 | 22.38 |
| Destination Spec. — color | ||||||||||
| Veo-3 | 63.8% | 39.3% | 47.9% | 37.5% | 1.08 | 0.41 | 7.40 | 3.65 | 1.11 | 16.17 |
| Sora-2 | 31.0% | 13.9% | 34.1% | 17.1% | 1.89 | 0.66 | 11.93 | 3.65 | 1.89 | 18.92 |
| Wan-2.2 | 37.9% | 19.3% | 36.4% | 21.8% | 2.93 | 0.85 | 17.64 | 3.65 | 1.84 | 26.49 |
| Destination Spec. — object | ||||||||||
| Veo-3 | 70.7% | 44.3% | 47.4% | 40.7% | 1.26 | 0.43 | 7.95 | 3.65 | 1.12 | 14.19 |
| Sora-2 | 50.0% | 23.1% | 39.9% | 26.2% | 1.65 | 0.52 | 10.42 | 3.65 | 1.50 | 12.12 |
| Wan-2.2 | 50.0% | 27.9% | 43.0% | 29.0% | 1.49 | 0.53 | 10.02 | 3.65 | 1.22 | 14.44 |
H.3 narrative: auto-metrics rate Veo-3 at 73.33% success while humans rate it at only 25.00%. In floor01 Veo-3 and Nano-banana both reach 73.33% Overall Success; in floor02plus Veo-3 degrades to 46.67% while Nano-banana holds 75.00%. Nano-banana scales with fidelity 60.00% (quality03) → 82.50% (quality05); Veo-3 plateaus at ≤62.50%.
T6 — Top-down View Navigation
Table 29: Quantitative results for the 2D Top-down Navigation benchmark (VLM-based)
Compares Sora-2, Veo-3, Nano-banana, GPT-4o-image (all values %). Columns: Success Score (2D) | Oracle Success Score (2D) | Object Semantic | Agent Consistency | Spatial Alignment | Destination Integrity | Scene Consistency | Physics Validness (gate) | Overall Success (holistic).
| Axis / Level | Model | Success Score (2D) | Oracle Success Score (2D) | Object Semantic | Agent Consistency | Spatial Alignment | Destination Integrity | Scene Consistency | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Env. Complexity — floor01 | Veo-3 | 65.71 | 85.71 | 77.14 | 62.86 | 91.43 | 65.71 | 45.71 | 54.29 | 37.14 |
| Sora-2 | 22.81 | 47.37 | 82.46 | 68.42 | 78.95 | 7.02 | 8.77 | 52.63 | 1.75 | |
| Nano-banana | 41.67 | 52.78 | 88.89 | 50.00 | 77.78 | 33.33 | 38.89 | 50.00 | 5.56 | |
| GPT-4o-image | 10.34 | 15.52 | 55.17 | 6.90 | 56.90 | 1.72 | 1.72 | 6.90 | 1.72 | |
| Env. Complexity — floor02plus | Veo-3 | 25.00 | 83.33 | 61.11 | 44.44 | 80.56 | 33.33 | 45.71 | 33.33 | 13.89 |
| Sora-2 | 28.81 | 55.93 | 89.83 | 54.24 | 86.44 | 15.25 | 23.73 | 49.15 | 5.08 | |
| Nano-banana | 38.89 | 52.78 | 86.11 | 55.56 | 86.11 | 36.11 | 66.67 | 50.00 | 16.67 | |
| GPT-4o-image | 16.67 | 23.33 | 51.67 | 11.67 | 41.67 | 5.00 | 10.00 | 10.00 | 5.00 | |
| View Fidelity — quality03 | Veo-3 | 45.83 | 79.17 | 66.67 | 33.33 | 75.00 | 50.00 | 43.48 | 29.17 | 20.83 |
| Sora-2 | 34.21 | 52.63 | 92.11 | 63.16 | 81.58 | 18.42 | 18.42 | 57.89 | 7.89 | |
| Nano-banana | 41.67 | 58.33 | 83.33 | 45.83 | 75.00 | 41.67 | 58.33 | 37.50 | 12.50 | |
| GPT-4o-image | 21.05 | 23.68 | 65.79 | 13.16 | 57.89 | 5.26 | 5.26 | 10.53 | 5.26 | |
| View Fidelity — quality04 | Veo-3 | 37.50 | 83.33 | 70.83 | 58.33 | 95.83 | 41.67 | 45.83 | 54.17 | 29.17 |
| Sora-2 | 23.08 | 46.15 | 87.18 | 66.67 | 89.74 | 7.69 | 17.95 | 51.28 | 2.56 | |
| Nano-banana | 37.50 | 50.00 | 87.50 | 50.00 | 83.33 | 33.33 | 37.50 | 50.00 | 4.17 | |
| GPT-4o-image | 12.50 | 17.50 | 52.50 | 12.50 | 50.00 | 5.00 | 10.00 | 12.50 | 5.00 | |
| View Fidelity — quality05 | Veo-3 | 52.17 | 91.30 | 69.57 | 69.57 | 86.96 | 56.52 | 47.83 | 47.83 | 26.09 |
| Sora-2 | 20.51 | 56.41 | 79.49 | 53.85 | 76.92 | 7.69 | 12.82 | 43.59 | 0.00 | |
| Nano-banana | 41.67 | 50.00 | 91.67 | 62.50 | 87.50 | 29.17 | 62.50 | 62.50 | 16.67 | |
| GPT-4o-image | 7.50 | 17.50 | 42.50 | 2.50 | 40.00 | 0.00 | 2.50 | 2.50 | 0.00 | |
| Trajectory Distance — noturn | Veo-3 | 36.11 | 83.33 | 61.11 | 52.78 | 83.33 | 38.89 | 54.29 | 38.89 | 25.00 |
| Sora-2 | 29.31 | 60.34 | 89.66 | 62.07 | 84.48 | 15.52 | 18.97 | 48.28 | 1.72 | |
| Nano-banana | 41.67 | 52.78 | 88.89 | 50.00 | 83.33 | 36.11 | 52.78 | 44.44 | 11.11 | |
| GPT-4o-image | 20.34 | 23.73 | 57.63 | 10.17 | 57.63 | 5.08 | 8.47 | 10.17 | 5.08 | |
| Trajectory Distance — oneturn | Veo-3 | 54.29 | 85.71 | 77.14 | 54.29 | 88.57 | 60.00 | 37.14 | 48.57 | 25.71 |
| Sora-2 | 22.41 | 43.10 | 82.76 | 60.34 | 81.03 | 6.90 | 13.79 | 53.45 | 5.17 | |
| Nano-banana | 38.89 | 52.78 | 86.11 | 55.56 | 80.56 | 33.33 | 52.78 | 55.56 | 11.11 | |
| GPT-4o-image | 6.78 | 15.25 | 49.15 | 8.47 | 40.68 | 1.69 | 3.39 | 6.78 | 1.69 | |
| Destination Spec — color mark | Veo-3 | 52.78 | 91.67 | 80.56 | 69.44 | 91.67 | 58.33 | 63.89 | 61.11 | 38.89 |
| Sora-2 | 17.24 | 53.45 | 77.59 | 56.90 | 77.59 | 12.07 | 24.14 | 44.83 | 1.72 | |
| Nano-banana | 58.33 | 77.78 | 91.67 | 44.44 | 86.11 | 55.56 | 55.56 | 41.67 | 13.89 | |
| GPT-4o-image | 10.34 | 13.79 | 39.66 | 3.45 | 34.48 | 1.72 | 3.45 | 3.45 | 1.72 | |
| Destination Spec — location description | Veo-3 | 37.14 | 77.14 | 57.14 | 37.14 | 80.00 | 40.00 | 26.47 | 25.71 | 11.43 |
| Sora-2 | 34.48 | 50.00 | 94.83 | 65.52 | 87.93 | 10.34 | 8.62 | 56.90 | 5.17 | |
| Nano-banana | 22.22 | 27.78 | 83.33 | 61.11 | 77.78 | 13.89 | 50.00 | 58.33 | 8.33 | |
| GPT-4o-image | 16.67 | 25.00 | 66.67 | 15.00 | 63.33 | 5.00 | 8.33 | 13.33 | 5.00 |
Table 30: Top-down View Real-World Navigation — View-Syn (higher is better except LPIPS)
| Model | Modality | FVD Norm ↑ | FID Norm ↑ | SSIM ↑ | PSNR ↑ | LPIPS ↓ |
|---|---|---|---|---|---|---|
| Sora-2 | Video | 0.7079 | – | 0.4072 | 7.21 | 0.8345 |
| Veo-3 | Video | 0.6731 | – | 0.1471 | 6.65 | 0.8800 |
| Wan-2.2 | Video | 0.6285 | – | 0.4333 | 7.81 | 0.7955 |
| GPT-image | Image | – | 0.6153 | – | – | – |
| Nano-banana | Image | – | 0.6153 | – | – | – |
| Qwen-image | Image | – | 0.6003 | – | – | – |
| GPT-image-1.5 | Image | – | 0.5247 | – | – | – |
| Nano-banana-pro | Image | – | 0.4890 | – | – | – |
G.4 narrative: Top-down is the hardest embodied row overall. Nano-banana Pro leads at 33.05%, GPT-image-1.5 second at 26.27%, Veo-3 third at 19.49%.
T7 — 3D Real-World Navigation
Table 31: Quantitative results for the 3D Real-world Navigation benchmark (VLM-based)
Compares Sora-2, Veo-3, Nano-banana, GPT-4o-image (all values %). Columns: S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success.
Environmental Complexity — floor01
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 85.00 | 91.67 | 81.67 | 86.67 | 63.33 | 53.33 | 15.00 | 11.67 | 40.00 | 3.33 |
| Sora-2 (Video) | 88.89 | 97.22 | 72.22 | 83.33 | 97.22 | 77.78 | 44.44 | 0.00 | 55.00 | 0.00 |
| Nano-banana (Image) | 75.00 | 75.00 | 100.00 | 97.22 | 97.22 | 75.00 | 86.11 | 75.00 | 97.22 | 75.00 |
| GPT-4o-image (Image) | 13.33 | 13.33 | 88.33 | 48.33 | 58.33 | 13.33 | 21.67 | 13.33 | 35.00 | 13.33 |
Environmental Complexity — floor02plus
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 68.33 | 78.33 | 75.00 | 68.33 | 73.33 | 38.33 | 18.33 | 5.00 | 33.33 | 0.00 |
| Sora-2 (Video) | 72.22 | 77.78 | 69.44 | 75.00 | 83.33 | 58.33 | 36.11 | 0.00 | 60.00 | 0.00 |
| Nano-banana (Image) | 83.33 | 86.11 | 94.44 | 97.22 | 94.44 | 86.11 | 100.00 | 83.33 | 94.44 | 83.33 |
| GPT-4o-image (Image) | 15.00 | 16.67 | 75.00 | 56.67 | 51.67 | 15.00 | 15.00 | 13.33 | 40.00 | 13.33 |
View Fidelity — quality03
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 65.00 | 77.50 | 72.50 | 75.00 | 70.00 | 40.00 | 15.00 | 5.00 | 30.00 | 0.00 |
| Sora-2 (Video) | 75.00 | 83.33 | 79.17 | 83.33 | 91.67 | 66.67 | 41.67 | 0.00 | 67.50 | 0.00 |
| Nano-banana (Image) | 75.00 | 79.17 | 95.83 | 95.83 | 91.67 | 79.17 | 91.67 | 75.00 | 91.67 | 75.00 |
| GPT-4o-image (Image) | 7.50 | 7.50 | 80.00 | 45.00 | 52.50 | 5.00 | 7.50 | 5.00 | 30.00 | 5.00 |
View Fidelity — quality04
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 82.50 | 85.00 | 75.00 | 85.00 | 65.00 | 50.00 | 20.00 | 10.00 | 42.50 | 2.50 |
| Sora-2 (Video) | 91.67 | 95.83 | 54.17 | 75.00 | 83.33 | 70.83 | 41.67 | 0.00 | 45.00 | 0.00 |
| Nano-banana (Image) | 87.50 | 87.50 | 100.00 | 100.00 | 100.00 | 87.50 | 91.67 | 87.50 | 100.00 | 87.50 |
| GPT-4o-image (Image) | 17.50 | 20.00 | 80.00 | 57.50 | 52.50 | 20.00 | 27.50 | 17.50 | 40.00 | 17.50 |
View Fidelity — quality05
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 82.50 | 92.50 | 87.50 | 72.50 | 70.00 | 47.50 | 15.00 | 10.00 | 37.50 | 2.50 |
| Sora-2 (Video) | 75.00 | 83.33 | 79.17 | 79.17 | 95.83 | 66.67 | 37.50 | 0.00 | 60.00 | 0.00 |
| Nano-banana (Image) | 75.00 | 75.00 | 95.83 | 95.83 | 95.83 | 75.00 | 95.83 | 75.00 | 95.83 | 75.00 |
| GPT-4o-image (Image) | 17.50 | 17.50 | 85.00 | 55.00 | 60.00 | 17.50 | 20.00 | 17.50 | 42.50 | 17.50 |
Trajectory Distance — short
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 73.33 | 85.00 | 81.67 | 78.33 | 68.33 | 46.67 | 16.67 | 8.33 | 38.33 | 3.33 |
| Sora-2 (Video) | 72.22 | 83.33 | 80.56 | 77.78 | 86.11 | 63.89 | 41.67 | 0.00 | 55.00 | 0.00 |
| Nano-banana (Image) | 77.78 | 80.56 | 97.22 | 97.22 | 94.44 | 80.56 | 88.89 | 77.78 | 94.44 | 77.78 |
| GPT-4o-image (Image) | 13.33 | 13.33 | 85.00 | 56.67 | 58.33 | 11.67 | 16.67 | 11.67 | 45.00 | 11.67 |
Trajectory Distance — long
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 80.00 | 85.00 | 75.00 | 76.67 | 68.33 | 45.00 | 16.67 | 8.33 | 35.00 | 0.00 |
| Sora-2 (Video) | 88.89 | 91.67 | 61.11 | 80.56 | 94.44 | 72.22 | 38.89 | 0.00 | 60.00 | 0.00 |
| Nano-banana (Image) | 80.56 | 80.56 | 97.22 | 97.22 | 97.22 | 80.56 | 97.22 | 80.56 | 97.22 | 80.56 |
| GPT-4o-image (Image) | 15.00 | 16.67 | 78.33 | 48.33 | 51.67 | 16.67 | 20.00 | 15.00 | 30.00 | 15.00 |
Destination Specification — color mark
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 93.33 | 96.67 | 66.67 | 73.33 | 70.00 | 31.67 | 28.33 | 15.00 | 28.33 | 3.33 |
| Sora-2 (Video) | 88.89 | 91.67 | 58.33 | 75.00 | 88.89 | 69.44 | 36.11 | 0.00 | 51.67 | 0.00 |
| Nano-banana (Image) | 80.56 | 80.56 | 100.00 | 100.00 | 100.00 | 80.56 | 91.67 | 80.56 | 100.00 | 80.56 |
| GPT-4o-image (Image) | 16.67 | 18.33 | 71.67 | 46.67 | 45.00 | 18.33 | 20.00 | 16.67 | 30.00 | 16.67 |
Destination Specification — location description
| Model | S.S.(3D) | O.S.(3D) | Obj.Sem. | Agent Con. | Spa.Ali. | Des.Inte. | Scene Con. | Success(3D) Orig.Dest. | Physics Validness | Overall Success |
|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 (Video) | 60.00 | 73.33 | 90.00 | 81.67 | 66.67 | 60.00 | 5.00 | 1.67 | 45.00 | 0.00 |
| Sora-2 (Video) | 72.22 | 83.33 | 83.33 | 83.33 | 91.67 | 66.67 | 44.44 | 0.00 | 63.33 | 0.00 |
| Nano-banana (Image) | 77.78 | 80.56 | 94.44 | 94.44 | 91.67 | 80.56 | 94.44 | 77.78 | 91.67 | 77.78 |
| GPT-4o-image (Image) | 11.67 | 11.67 | 91.67 | 58.33 | 65.00 | 10.00 | 16.67 | 10.00 | 45.00 | 10.00 |
Table 32: 3D Real-World Navigation — View-Syn (higher is better except LPIPS)
| Model | Modality | FVD Norm ↑ | FID Norm ↑ | SSIM ↑ | PSNR ↑ | LPIPS ↓ |
|---|---|---|---|---|---|---|
| Veo-3 | Video | 0.7708 | – | 0.2987 | 7.35 | 0.8205 |
| Wan-2.2 | Video | 0.7107 | – | 0.4562 | 9.11 | 0.7652 |
| Sora-2 | Video | 0.6870 | – | 0.4471 | 9.19 | 0.8052 |
| GPT-image-1.5 | Image | – | 0.6747 | – | – | – |
| Nano-banana-pro | Image | – | 0.6740 | – | – | – |
| GPT-image | Image | – | 0.5673 | – | – | – |
| Nano-banana | Image | – | 0.5673 | – | – | – |
| Qwen-image | Image | – | 0.5627 | – | – | – |
G.4.1: Nano-banana Pro reaches 85.00%, followed by Nano-banana (79.17%) and GPT-image-1.5 (77.92%). Video models remain below 25%.
T8 — SLAG
Table 33: Quantitative results for the SLAG benchmark (VLM-based)
Compares Sora-2, Veo-3, Nano-banana, GPT-4o-image (all values %). Column key: SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall.
Environmental Complexity — floor01
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 52.54% | 44.07% | 59.32% | 49.15% | 44.07% | 72.88% | 52.54% | 67.80% | 20.34% | 45.76% | 18.64% | 45.76% | 11.86% |
| Sora-2 | 29.31% | 27.59% | 34.48% | 37.93% | 55.17% | 84.48% | 84.48% | 72.41% | 32.76% | 79.31% | 22.41% | 65.52% | 10.34% |
| Nano-banana | 55.56% | 41.67% | 55.56% | 47.22% | 88.89% | 83.33% | 83.33% | 77.78% | 50.00% | 100.00% | 38.89% | 69.44% | 27.78% |
| GPT-4o-image | 35.00% | 28.33% | 36.67% | 31.67% | 66.67% | 81.67% | 76.67% | 61.67% | 26.67% | 98.33% | 25.00% | 58.33% | 25.00% |
Environmental Complexity — floor02plus
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 21.05% | 38.60% | 26.32% | 42.11% | 33.33% | 50.88% | 42.11% | 52.63% | 31.58% | 57.89% | 15.79% | 31.58% | 10.53% |
| Sora-2 | 29.31% | 32.76% | 31.03% | 37.93% | 48.28% | 82.76% | 67.24% | 68.97% | 44.83% | 89.66% | 20.69% | 56.90% | 15.52% |
| Nano-banana | 40.00% | 56.67% | 40.00% | 63.33% | 76.67% | 86.67% | 86.67% | 70.00% | 50.00% | 96.67% | 36.67% | 60.00% | 30.00% |
| GPT-4o-image | 12.07% | 17.24% | 13.79% | 20.69% | 43.10% | 68.97% | 48.28% | 37.93% | 17.24% | 94.83% | 6.90% | 34.48% | 6.90% |
View Fidelity — quality03
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 35.00% | 40.00% | 35.00% | 40.00% | 35.00% | 60.00% | 45.00% | 52.50% | 22.50% | 45.00% | 15.00% | 37.50% | 10.00% |
| Sora-2 | 27.50% | 27.50% | 27.50% | 32.50% | 47.50% | 85.00% | 77.50% | 75.00% | 37.50% | 77.50% | 22.50% | 62.50% | 17.50% |
| Nano-banana | 50.00% | 62.50% | 50.00% | 62.50% | 91.67% | 91.67% | 95.83% | 91.67% | 58.33% | 100.00% | 41.67% | 83.33% | 41.67% |
| GPT-4o-image | 30.00% | 25.00% | 32.50% | 32.50% | 65.00% | 75.00% | 65.00% | 52.50% | 20.00% | 97.50% | 17.50% | 52.50% | 17.50% |
View Fidelity — quality04
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 46.15% | 51.28% | 51.28% | 58.97% | 43.59% | 66.67% | 46.15% | 69.23% | 33.33% | 48.72% | 25.64% | 35.90% | 15.38% |
| Sora-2 | 26.32% | 34.21% | 26.32% | 42.11% | 55.26% | 84.21% | 71.05% | 71.05% | 34.21% | 84.21% | 21.05% | 60.53% | 10.53% |
| Nano-banana | 45.83% | 45.83% | 45.83% | 54.17% | 70.83% | 87.50% | 79.17% | 58.33% | 50.00% | 95.83% | 37.50% | 54.17% | 20.83% |
| GPT-4o-image | 25.00% | 30.00% | 25.00% | 30.00% | 57.50% | 72.50% | 62.50% | 50.00% | 32.50% | 97.50% | 22.50% | 47.50% | 22.50% |
View Fidelity — quality05
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 29.73% | 32.43% | 43.24% | 37.84% | 37.84% | 59.46% | 51.35% | 59.46% | 21.62% | 62.16% | 10.81% | 43.24% | 8.11% |
| Sora-2 | 34.21% | 28.95% | 44.74% | 39.47% | 52.63% | 81.58% | 78.95% | 65.79% | 44.74% | 92.11% | 21.05% | 60.53% | 10.53% |
| Nano-banana | 50.00% | 33.33% | 50.00% | 44.44% | 88.89% | 72.22% | 77.78% | 72.22% | 38.89% | 100.00% | 33.33% | 55.56% | 22.22% |
| GPT-4o-image | 15.79% | 13.16% | 18.42% | 15.79% | 42.11% | 78.95% | 60.53% | 47.37% | 13.16% | 94.74% | 7.89% | 39.47% | 7.89% |
Trajectory Distance — short
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 38.98% | 40.68% | 45.76% | 45.76% | 44.07% | 61.02% | 54.24% | 64.41% | 28.81% | 62.71% | 20.34% | 42.37% | 15.25% |
| Sora-2 | 40.35% | 31.58% | 43.86% | 40.35% | 52.63% | 85.96% | 82.46% | 75.44% | 47.37% | 89.47% | 29.82% | 66.67% | 15.79% |
| Nano-banana | 54.55% | 51.52% | 54.55% | 54.55% | 87.88% | 90.91% | 90.91% | 90.91% | 57.58% | 100.00% | 42.42% | 78.79% | 36.36% |
| GPT-4o-image | 27.12% | 20.34% | 28.81% | 25.42% | 55.93% | 77.97% | 62.71% | 50.85% | 28.81% | 96.61% | 18.64% | 45.76% | 18.64% |
Trajectory Distance — long
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 35.09% | 42.11% | 40.35% | 45.61% | 33.33% | 63.16% | 40.35% | 56.14% | 22.81% | 40.35% | 14.04% | 35.09% | 7.02% |
| Sora-2 | 18.64% | 28.81% | 22.03% | 35.59% | 50.85% | 81.36% | 69.49% | 66.10% | 30.51% | 79.66% | 13.56% | 55.93% | 10.17% |
| Nano-banana | 42.42% | 45.45% | 42.42% | 54.55% | 78.79% | 78.79% | 78.79% | 57.58% | 42.42% | 96.97% | 33.33% | 51.52% | 21.21% |
| GPT-4o-image | 20.34% | 25.42% | 22.03% | 27.12% | 54.24% | 72.88% | 62.71% | 49.15% | 15.25% | 96.61% | 13.56% | 47.46% | 13.56% |
Destination Specification — color mark
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 53.45% | 51.72% | 60.34% | 56.90% | 43.10% | 56.90% | 46.55% | 65.52% | 32.76% | 67.24% | 25.86% | 34.48% | 17.24% |
| Sora-2 | 44.64% | 46.43% | 51.79% | 57.14% | 48.21% | 82.14% | 73.21% | 67.86% | 62.50% | 92.86% | 33.93% | 58.93% | 19.64% |
| Nano-banana | 81.25% | 59.38% | 81.25% | 68.75% | 87.50% | 84.38% | 87.50% | 78.12% | 71.88% | 100.00% | 62.50% | 71.88% | 50.00% |
| GPT-4o-image | 44.83% | 36.21% | 44.83% | 41.38% | 55.17% | 70.69% | 62.07% | 51.72% | 37.93% | 96.55% | 31.03% | 46.55% | 31.03% |
Destination Specification — location description
| Model | SS3D | SS2D | OS3D | OS2D | TrajAlign | ObjSem | AgentCon | SpaAlign | DestInteg | SceneCon | Succ3D-OrigDest | PhysValid | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Veo-3 | 20.69% | 31.03% | 25.86% | 34.48% | 34.48% | 67.24% | 48.28% | 55.17% | 18.97% | 36.21% | 8.62% | 43.10% | 5.17% |
| Sora-2 | 15.00% | 15.00% | 15.00% | 20.00% | 55.00% | 85.00% | 78.33% | 73.33% | 16.67% | 76.67% | 10.00% | 63.33% | 6.67% |
| Nano-banana | 17.65% | 38.24% | 17.65% | 41.18% | 79.41% | 85.29% | 82.35% | 70.59% | 29.41% | 97.06% | 14.71% | 58.82% | 8.82% |
| GPT-4o-image | 3.33% | 10.00% | 6.67% | 11.67% | 55.00% | 80.00% | 63.33% | 48.33% | 6.67% | 96.67% | 1.67% | 46.67% | 1.67% |
Table 34 / K.3.2 View-Syn (SLAG)
The K.3.2 "View-Syn Evaluation" subsection header exists but contains no body text and no results table in the paper. View-Syn numeric results for SLAG: not specified in paper. (The digest also notes a Table 34 of FVD/FID/SSIM/PSNR/LPIPS is associated with SLAG but no numeric values were transcribed.)
G.4.1: Nano-banana Pro leads SLAG at 37.29%, followed by GPT-image-1.5 (31.36%) and Nano-banana (28.79%).
T9 & T10 — Physical Commonsense
Table 35: Distribution of Physical Commonsense evaluation samples
| Task Axis | Total Samples |
|---|---|
| Physical Concepts (Atomic Interactions) | 25 |
| Sports Scenarios (Compositional Contexts) | 25 |
| Total | 50 |
Table 36: Quantitative results for the Physical Commonsense task
Models: Veo-3, Sora-2, Wan-2.2 (video generative models), evaluated with Gemini-2.5-Pro. Higher is better (↑) for all.
| Model | Physics Accuracy ↑ | Motion Quality ↑ | Visual Realism ↑ | Prompt Adherence ↑ | Overall ↑ |
|---|---|---|---|---|---|
| Scenario Type: Physical Concepts | |||||
| Veo-3 | 62.50% | 54.17% | 83.33% | 50.00% | 41.67% |
| Sora-2 | 84.00% | 80.00% | 96.00% | 76.00% | 76.00% |
| Wan-2.2 | 58.67% | 53.33% | 72.00% | 38.67% | 26.67% |
| Scenario Type: Sports Scenarios | |||||
| Veo-3 | 80.00% | 68.00% | 92.00% | 68.00% | 60.00% |
| Sora-2 | 88.00% | 72.00% | 88.00% | 68.00% | 64.00% |
| Wan-2.2 | 42.67% | 33.33% | 96.00% | 21.33% | 21.33% |
| Average | |||||
| Veo-3 | 71.43% | 61.22% | 87.76% | 59.18% | 51.02% |
| Sora-2 | 86.00% | 76.00% | 92.00% | 72.00% | 70.00% |
| Wan-2.2 | 50.67% | 43.33% | 84.00% | 30.00% | 24.00% |
Note: the two task-level rows reported in Table 4 are Physical Concepts and Sports separately; the "Average" row here is a diagnostic aggregate.
Table 37: Fine-grained VLM-based evaluation results by Sports Scenarios attributes
Overall success rate (%) across sport types (Ballet, Diving, Skiing, Swimming) and difficulty levels (Easy, Medium, Hard).
| Model | Ballet | Diving | Skiing | Swimming | Easy | Medium | Hard |
|---|---|---|---|---|---|---|---|
| Veo-3 | 33.3% | 50.0% | 71.4% | 83.3% | 60.0% | 62.5% | 57.1% |
| Sora-2 | 33.3% | 50.0% | 85.7% | 83.3% | 60.0% | 75.0% | 57.1% |
| Wan-2.2 | 44.4% | 0.0% | 28.6% | 11.1% | 16.7% | 33.3% | 14.3% |
Table 38: Fine-grained VLM-based evaluation results by Physical Concepts attributes
Overall success rate (%) across states-of-matter interaction types and difficulty levels.
| Model | Solid-Solid | Solid-Fluid | Fluid-Fluid | Action/Other | Easy | Hard |
|---|---|---|---|---|---|---|
| Veo-3 | 0.0% | 75.0% | 50.0% | 40.0% | 53.3% | 22.2% |
| Sora-2 | 100.0% | 75.0% | 100.0% | 66.7% | 75.0% | 77.8% |
| Wan-2.2 | 33.3% | 25.0% | 83.3% | 17.8% | 35.4% | 11.1% |
Key findings (L.5): Sora-2 leads both final task rows — 76.00% on Physical Concepts and 64.00% on Sports (average 70.00%), ahead of Veo-3 (51.02%) and Wan-2.2 (24.00%). Wan-2.2 produces highly photorealistic videos (84–96% Visual Realism) yet fails prompts/physics (24% aggregate Overall). Solid-Solid interactions are most challenging (Veo-3 = 0% on rigid body collisions vs 75% on Solid-Fluid).
Evaluated models, sampling, totals, and cost
Evaluated models (exact names, versions, sizes)
Video generators (3):
- Veo-3 — closed-source, API; parameter count not disclosed.
- Sora-2 — closed-source, API; parameter count not disclosed.
- Wan-2.2 — open-source; used in its image-to-video I2V-A14B Mixture-of-Experts configuration, 27B total parameters with about 14B active per denoising step.
Image generators (5):
- Nano-banana — closed-source, API.
- Nano-banana Pro — closed-source, API.
- GPT-4o-image — closed-source, API.
- GPT-image-1.5 — closed-source, API.
- Qwen-image / Qwen-Image — open-source; 20B-parameter Multimodal Diffusion Transformer generating at 1024×1024 resolution.
LLM/VLM text-solution baselines (2, abstract-reasoning tasks only):
- Gemini-3-Flash — closed-source, API.
- Gemini-3-Pro — closed-source, API.
- The Gemini text baselines produce generations only for Maze, Sudoku, and Math (Section A.1).
Judge / evaluator model (all tasks): Gemini-2.5-Pro (VLM judge for tasks without reliable pixel-/symbol-level verification). Deterministic evaluators: Maze pixel-based; Sudoku OCR-based (PaddleOCR).
Samples per prompt and totals
- Five (5) outputs per prompt from each generative model (default settings for closed models, recommended settings for open ones; no task-specific tuning).
- Total evaluation samples (task instances): 1,853.
- Total generations (5 samples per instance):
- A video model produces 9,265 generations.
- An image model produces 9,015 generations (the eight non-physics tasks).
- A Gemini text baseline produces 4,335 generations (Maze, Sudoku, and Math).
Compute (Section A)
- On a single A100 80GB GPU: Qwen-Image renders one image in ≈40 s; Wan-2.2 produces one clip in ≈520 s.
- Open-source GPU consumption: Qwen-Image ≈100 A100 GPU-hours; Wan-2.2 ≈1,340 A100 GPU-hours; ≈1,440 GPU-hours total.
- Assumed rental rate: ≈$1.5 per A100 GPU-hour.
- Assumed API prices: ≈$0.05 per image; $0.40 and $0.10 per second of generated video for Veo-3 and Sora-2 respectively; plus metered token usage for the Gemini baselines and the Gemini-2.5-Pro judge.
- Deterministic Maze and Sudoku checks run on CPU at negligible cost.
Table 6: Per-model cost breakdown (Est. cost USD)
| Model / group | Est. cost (USD) |
|---|---|
| Veo-3 (API) | ~30k |
| Sora-2 (API) | ~9k |
| Four closed image models (API) — Nano-banana, Nano-banana Pro, GPT-4o-image, GPT-image-1.5 | ~1.8k |
| Gemini baselines + judge (API) | ~3k |
| Wan-2.2 (A100, ~1,340 h) | ~2.0k |
| Qwen-Image (A100, ~100 h) | ~0.15k |
| Total | ~46k |
Estimated total inference cost of the benchmark: ~$46k, dominated by video generation.
URLs
- Project page: https://zefan-cai.github.io/MMGR
- Code: https://github.com/Zefan-Cai/MMGR
- Data: https://huggingface.co/datasets/ZefanCai/MMGR
- Math source (MAA): https://maa.org/
- ARC sources: https://github.com/fchollet/ARC , https://github.com/michaelhodel/re-arc , https://github.com/google/ARC-GEN