3. Start judge server (separate terminal)

July 7, 2026 · View on GitHub

ByteDance Seed

EdgeBench


Project Tech Report Dataset Docs WeChat Group Discord


Overview

EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks along with the full evaluation framework.

Analyzing ~38,000 hours of agent interaction on all 134 tasks, we find that performance follows a log-sigmoid scaling law as a function of interaction time (R2=0.998R^2 = 0.998). See the tech report for details.

Log-sigmoid scaling fit across 134 tasks

Leaderboard

Full Benchmark (134 tasks)

Model@2h@4h@6h@8h@10h@12h
Claude Opus 4.839.045.748.149.850.951.3
GPT-5.536.842.144.546.347.648.4
GPT-5.429.734.036.538.038.939.3
GLM-5.126.030.432.934.936.537.4
DS-V4-Pro23.327.129.029.930.931.0
Category Scores @12h (134 tasks)
ModelScientific & MLSystems & SEOptimizationKnowledgeFormalGames
Claude Opus 4.848.567.436.547.055.039.3
GPT-5.544.365.033.645.750.039.1
GPT-5.433.554.127.938.840.829.0
GLM-5.133.850.926.443.524.629.3
DS-V4-Pro30.043.021.537.014.116.9

Open-Source Subset (51 tasks)

Model@2h@4h@6h@8h@10h@12h
Claude Opus 4.833.238.540.842.143.344.2
GPT-5.531.236.038.240.342.143.1
GPT-5.425.028.230.332.133.334.2
GLM-5.121.424.226.828.229.130.4
DS-V4-Pro17.121.122.923.825.125.7
Category Scores @12h (51 tasks)
ModelScientific & MLSystems & SEOptimizationKnowledgeFormalGames
Claude Opus 4.838.962.038.238.740.939.3
GPT-5.533.260.532.338.449.039.1
GPT-5.424.650.129.931.630.229.0
GLM-5.126.843.626.731.019.929.3
DS-V4-Pro31.137.624.133.212.716.9
Per-Task Scores by Time Budget (51 tasks)

Each model cell reports scores at @2h / @4h / @6h / @8h / @10h / @12h. Missing valid results are shown as .

TaskCategoryOpus 4.8GPT-5.5GPT-5.4GLM-5.1DS-V4-Pro
bipedalwalker_locomotion_rlScientific & ML16.7/20.8/22.4/23.3/23.3/23.314.7/14.9/15.2/15.2/16.0/21.013.9/13.9/13.9/14.5/14.5/17.513.9/20.3/21.5/22.5/22.5/22.58.9/14.8/17.6/20.4/20.4/20.6
borden_source_inversionScientific & ML7.5/19.8/26.2/28.5/38.5/48.420.1/27.0/29.4/37.8/38.1/38.57.2/7.3/7.6/7.9/8.0/8.07.0/10.3/12.0/12.3/12.5/15.17.0/11.6/15.1/26.6/36.7/38.2
dabic_gravity_inversionScientific & ML9.5/15.2/15.7/17.4/17.5/17.515.9/16.2/16.7/17.0/17.2/17.314.6/14.6/15.5/15.5/15.0/15.09.2/13.7/16.0/16.5/16.5/17.1—/12.7/12.7/12.7/13.0/13.8
graph_node_classificationScientific & ML59.4/62.7/65.0/65.6/66.5/66.654.7/55.1/55.1/55.3/55.9/56.054.9/56.2/56.5/56.9/57.5/57.649.4/52.3/52.3/52.3/52.3/52.346.0/48.2/49.2/51.3/51.7/51.8
ann_vector_search_qpsSystems & SE26.2/57.0/58.6/58.7/59.4/59.722.3/34.3/35.1/36.0/40.0/40.727.5/30.2/44.5/45.2/49.7/50.26.7/24.4/25.6/25.6/26.1/38.39.4/19.6/22.4/22.8/23.8/23.8
arc_compiler_runtimeSystems & SE49.3/52.0/52.0/52.0/52.0/52.055.5/56.5/60.9/70.3/71.0/72.445.1/46.5/49.8/49.8/50.0/50.047.7/48.0/48.4/48.7/48.7/48.740.3/41.7/44.2/44.2/44.2/44.2
exchange_core_throughputSystems & SE40.7/57.0/58.5/58.9/59.7/59.715.4/37.2/39.9/44.3/51.3/53.214.3/40.8/41.0/45.2/46.4/47.329.2/43.7/46.5/48.6/50.3/52.632.9/33.8/45.0/47.7/48.4/48.6
ffmpeg_swscale_reimplementationSystems & SE9.9/17.6/19.8/20.9/21.1/21.18.8/14.3/15.1/15.3/15.3/15.35.4/8.5/9.4/11.6/13.3/13.90.3/0.3/0.4/2.2/2.2/2.20.1/1.9/2.0/3.8/3.8/3.8
git_rewrite_in_zigSystems & SE22.0/22.8/22.8/22.8/23.1/23.116.1/16.9/17.7/18.2/18.2/18.49.6/13.8/14.0/14.2/14.2/15.412.0/20.2/23.3/23.4/23.4/23.58.5/13.5/16.0/17.6/17.8/17.9
integer_compression_codecSystems & SE69.4/69.7/74.8/74.9/75.2/75.361.1/67.6/73.9/73.9/74.3/74.438.6/40.9/41.2/42.2/42.2/42.323.5/27.3/28.5/28.7/28.9/28.915.9/16.0/16.2/16.2/16.2/16.2
juliet_vulnerability_analyzerSystems & SE71.9/74.9/75.4/75.6/75.6/75.681.0/83.2/85.4/86.8/87.4/89.852.9/66.1/74.3/76.0/76.8/77.259.3/60.7/62.8/63.5/63.5/63.546.8/63.1/66.1/66.2/66.2/66.2
rust_multicrate_reconstructionSystems & SE—/—/—/—/—/—27.5/42.6/53.1/54.9/57.8/57.816.7/19.9/21.3/21.4/21.4/21.424.8/24.8/25.2/25.2/37.5/38.520.5/21.7/22.7/23.1/23.5/23.6
schemathesis_config_modernizationSystems & SE82.5/85.0/86.1/87.4/87.4/87.779.1/82.2/82.9/83.2/83.6/84.067.2/68.8/68.8/71.7/71.7/71.958.3/59.7/60.4/61.2/61.7/61.754.3/54.3/55.3/55.3/55.3/55.6
schemathesis_datagen_pipelineSystems & SE68.0/70.2/70.2/70.2/70.2/70.254.6/54.6/56.7/56.7/56.7/56.756.6/56.6/56.6/56.6/56.6/56.662.1/64.2/64.2/67.0/67.0/67.047.9/50.1/52.3/52.3/52.3/52.3
schemathesis_reporting_observabilitySystems & SE73.9/75.6/76.2/76.2/76.2/76.276.6/76.6/76.6/76.6/77.1/77.170.0/73.7/74.7/75.7/76.2/76.261.9/61.9/61.9/61.9/61.9/61.959.4/62.4/63.0/63.0/65.0/65.0
vliw_kernel_optimizationSystems & SE74.0/76.0/77.7/79.5/79.6/80.971.6/75.7/77.1/79.5/83.1/85.675.7/77.0/77.2/78.7/79.1/79.15.6/9.5/27.5/35.0/35.9/35.90.2/24.9/28.1/33.0/33.9/34.1
ad_placement_optimizationOptimization65.2/66.1/66.9/67.1/67.4/67.744.0/53.3/59.5/61.6/62.9/62.941.8/42.4/43.1/47.7/47.9/48.148.7/52.7/53.3/56.5/58.5/58.825.5/28.5/35.2/35.8/36.2/36.2
apple_incremental_gameOptimization42.7/44.9/45.9/48.6/49.9/50.626.6/29.8/30.6/32.7/33.1/33.628.3/30.3/32.0/33.3/33.9/34.919.0/19.0/19.1/19.1/19.1/19.119.6/19.7/19.7/19.7/19.7/19.7
equivalence_class_divide_and_conquerOptimization11.2/15.3/17.0/20.1/20.8/21.311.8/15.5/15.8/21.3/22.2/22.414.5/17.0/18.3/18.7/20.2/20.33.8/4.2/10.0/8.0/10.6/10.60.7/1.8/3.2/3.2/3.4/3.4
grid_turing_robotOptimization34.7/37.1/37.3/39.6/40.3/40.340.4/41.6/41.9/42.0/42.1/42.226.8/26.8/27.2/28.9/28.9/28.920.0/21.0/24.6/24.6/24.6/25.723.7/24.1/24.1/24.1/24.2/24.2
jagua_nesting_optimizationOptimization11.2/17.8/24.5/31.4/41.0/44.215.9/19.4/20.0/20.6/21.3/21.622.4/23.0/23.9/24.0/24.1/24.18.9/9.0/10.0/12.2/12.3/12.410.7/20.2/23.7/26.7/28.1/28.4
molecular_self_assemblyOptimization22.4/33.4/34.0/34.1/34.4/34.720.2/20.3/20.5/20.7/20.7/20.720.8/21.1/21.1/21.5/21.5/21.610.0/12.5/12.9/13.0/13.1/13.219.4/21.7/21.8/21.8/21.9/21.9
order_addition_permutation_optimizationOptimization22.6/31.6/34.0/34.4/35.7/36.416.7/20.5/21.5/22.4/23.0/23.31.6/10.6/13.1/14.0/14.2/14.32.0/2.1/23.6/25.8/25.8/33.24.6/16.5/17.8/22.9/25.4/30.8
smt_solverOptimization10.3/17.4/19.0/23.1/23.3/23.97.2/7.8/8.4/8.6/8.6/8.66.7/7.9/8.9/9.1/9.1/9.22.7/2.7/2.7/2.7/2.7/3.61.4/2.8/3.3/3.3/3.3/3.3
treant_forestOptimization14.5/15.9/16.1/16.2/16.4/18.012.1/14.2/14.9/15.2/15.5/15.612.2/12.2/12.7/13.0/13.2/13.38.0/11.6/11.7/14.1/14.5/16.96.8/8.1/9.7/10.1/12.7/13.5
tree_block_partitioningOptimization21.5/30.1/32.4/36.8/37.7/37.728.8/31.1/33.0/33.0/35.0/36.423.1/26.8/28.8/32.9/34.3/34.312.1/15.4/17.1/19.3/20.3/23.411.2/11.8/11.9/11.9/14.6/16.1
triangulation_coloring_optimizationOptimization70.8/71.4/71.9/73.2/73.3/73.473.7/74.3/74.5/75.0/75.1/75.274.1/74.2/74.3/74.3/74.3/74.368.8/71.2/71.6/72.0/72.7/73.056.1/58.0/59.0/59.1/59.1/59.3
vehicle_routing_time_windowsOptimization72.5/72.6/72.9/73.6/73.7/74.088.7/89.0/89.4/89.7/89.7/90.885.3/88.6/88.7/89.5/89.5/89.676.6/76.6/76.6/76.6/76.6/77.954.7/76.8/81.9/82.2/82.9/83.1
vibrating_path_graph_coloringOptimization19.7/21.1/21.4/22.5/24.5/25.310.1/10.5/10.6/10.7/10.7/11.418.1/19.4/19.8/23.4/23.6/24.19.6/18.3/20.3/22.9/22.9/22.912.4/14.4/19.3/19.4/21.8/22.1
warehouse_forklift_routingOptimization7.7/9.5/10.4/10.5/11.1/11.29.8/11.0/11.8/11.9/12.1/12.60.0/0.0/0.0/0.0/0.0/0.0—/0.0/0.0/0.6/0.7/0.50.0/0.0/0.0/0.0/0.0/0.0
wireless_electricity_layoutOptimization6.5/13.7/14.4/14.5/14.5/14.56.2/6.9/7.1/7.1/7.1/7.210.9/11.1/11.1/11.1/11.1/11.17.2/9.4/6.6/8.1/9.4/9.50.0/0.0/0.0/0.0/0.0/0.0
college_english_exam_bankKnowledge24.8/28.3/34.8/35.5/35.8/39.824.5/35.5/35.5/35.5/37.8/37.830.7/30.7/31.3/34.0/34.0/34.522.2/26.0/29.3/30.0/32.3/32.519.2/21.7/22.5/22.7/29.2/34.7
cta_risk_budget_optimizationKnowledge42.7/44.8/45.3/45.3/45.3/46.143.8/45.8/46.7/46.7/46.7/46.746.0/49.0/49.0/49.0/49.8/49.838.1/44.8/49.0/49.6/49.6/49.644.0/45.6/46.9/46.9/48.1/48.1
k12_math_recommendationKnowledge23.6/38.5/41.4/42.0/43.7/44.338.5/42.4/42.9/43.5/43.9/44.025.9/29.0/30.0/30.8/31.1/31.424.8/25.7/31.9/32.5/32.7/32.725.6/26.3/26.8/25.7/26.0/26.3
portfolio_risk_calibrationKnowledge20.1/21.6/23.0/23.6/23.6/24.517.3/21.3/22.7/23.5/24.4/25.06.0/9.6/10.7/10.7/10.7/10.70.0/8.4/8.5/8.9/9.2/9.410.4/16.3/16.6/16.7/23.7/23.7
carleson_formalizationFormal4.3/7.7/11.0/12.7/15.0/16.86.0/9.5/13.2/16.5/25.3/26.51.8/3.5/4.6/5.4/6.3/7.11.0/1.7/2.0/2.2/2.2/2.20.8/1.3/2.0/2.0/2.3/2.5
combinatorial_games_formalizationFormal14.5/23.2/27.6/32.1/34.5/35.512.0/18.8/24.6/27.2/33.4/38.25.9/8.3/11.5/13.5/16.3/17.86.7/9.8/14.3/14.9/16.2/16.24.3/6.7/7.3/7.4/7.7/7.8
flt_regular_formalizationFormal31.0/41.8/50.6/50.6/50.6/50.643.7/48.3/50.6/66.7/75.1/75.11.5/19.5/28.4/41.8/46.0/48.314.4/13.4/16.5/18.8/18.8/38.75.7/11.9/14.6/14.9/17.2/17.6
lean_analysis_proofsFormal17.9/25.1/28.6/30.2/32.6/33.016.8/23.2/28.4/33.9/39.0/42.53.6/8.1/10.8/12.9/15.5/16.45.2/5.9/5.9/5.9/5.9/5.95.8/7.3/8.2/8.8/9.3/9.5
new_foundations_consistencyFormal28.9/36.2/50.0/62.7/64.2/65.113.7/38.2/55.1/56.4/65.1/66.53.3/12.2/14.9/20.5/30.7/39.82.2/3.3/5.1/21.9/24.6/27.02.2/3.4/6.5/7.2/10.5/11.4
ordinal_notation_well_foundednessFormal10.6/18.4/24.7/24.7/24.7/24.713.7/24.7/24.7/24.7/24.7/24.71.2/5.5/13.7/15.3/18.4/21.62.0/3.5/5.1/5.9/5.9/5.93.5/3.5/4.7/4.7/4.7/4.7
pfr_formalizationFormal32.4/36.9/38.8/40.2/45.6/46.330.7/38.3/41.9/47.6/52.7/60.010.2/13.7/27.1/34.0/35.9/38.98.3/14.9/22.5/26.5/31.3/33.59.9/14.7/16.5/17.8/18.5/19.1
sphere_eversion_formalizationFormal41.7/47.4/49.1/50.4/54.1/55.445.0/51.1/55.0/55.9/56.9/58.513.3/20.2/32.7/43.7/50.4/51.415.5/24.2/26.7/28.7/30.2/30.22.9/14.1/22.3/24.9/28.6/29.3
anchorhead_text_adventureGames13.3/19.3/19.7/20.3/22.3/22.315.0/26.3/31.7/34.3/35.3/36.35.0/11.7/13.0/13.3/14.7/17.710.7/17.3/19.7/20.3/20.3/20.32.0/6.0/7.3/8.0/12.3/14.7
dcss_dungeon_aiGames4.2/4.9/5.9/6.3/6.3/8.38.9/9.7/10.0/10.0/13.3/13.42.6/5.6/5.6/5.6/6.1/6.12.8/3.0/3.3/3.3/5.1/7.62.8/3.6/4.4/4.5/5.1/5.7
nethack_dungeon_agentGames29.7/35.3/36.7/37.3/41.9/41.916.6/17.6/18.1/20.6/21.3/22.510.9/14.1/15.2/15.8/17.0/20.42.3/2.3/15.3/21.6/21.6/21.61.0/1.4/2.9/3.2/3.2/3.3
openrct2_theme_park_aiGames24.4/24.4/26.0/26.0/27.5/27.528.5/28.6/32.7/37.3/37.4/37.623.0/23.1/23.1/23.1/23.1/23.135.1/36.2/36.2/36.2/36.2/36.224.4/24.4/24.4/26.0/26.0/26.0
openttd_transport_aiGames50.0/50.4/50.6/51.7/51.8/52.010.1/11.6/13.2/21.9/25.6/28.110.8/11.4/11.6/11.9/11.9/11.90.0/0.0/0.0/0.0/0.0/0.04.8/9.2/9.3/12.3/15.2/15.2
trinity_text_adventureGames25.0/28.0/29.3/30.0/30.0/30.022.3/26.7/28.7/36.3/36.3/40.016.3/19.7/22.7/23.3/23.7/27.016.0/18.7/24.3/26.0/26.7/26.716.3/16.3/17.7/20.0/20.0/20.3
tryst_text_adventureGames18.1/33.8/36.7/40.0/40.0/44.332.1/42.4/44.3/48.6/55.2/55.719.5/20.0/20.0/31.0/38.6/44.318.6/28.6/36.2/40.5/42.9/43.38.6/11.4/11.4/11.4/11.4/13.8
wesnoth_tactical_aiGames84.0/85.3/87.7/87.7/87.7/88.064.7/73.0/76.3/78.0/78.3/79.379.7/79.7/80.3/80.3/81.3/81.375.7/78.3/78.3/78.3/78.3/78.317.0/36.3/36.3/36.3/36.3/36.3

Task Taxonomy

EdgeBench contains 134 realistic, diverse tasks spanning six capability categories, of which 51 are publicly released. Each task is designed as a day-scale challenge with a performance ceiling high enough that no current agent can saturate it. Recorded human expert effort averages 57.2 hours per task (up to 320 hours).

EdgeBench Task Taxonomy

Evaluation Harness: SForge

EdgeBench is powered by SForge, a two-container evaluation harness built for long-horizon agent evaluation. Each task materializes as isolated work and judge Docker images — the agent only sees the work environment, while hidden tests run in ephemeral judge containers.

Key mechanisms:

  • Two-container isolation — work and judge environments are fully separated, preventing evaluation hacking at its root
  • Iterative evaluation with feedback — agents don't submit once at the end for a one-shot score; instead they submit throughout the run, receive granular feedback (pass rates, failing tests, scores), and improve in a closed loop until timeout — the best result across all submissions is the final score
  • Long-horizon execution — stop hooks prevent premature agent exit, auto-resume recovers from transient failures, and the Kubernetes backend enables parallel runs at scale

Quick Start

# Install (requires Docker Engine running on a Linux host)
pip install sforge

# 1. Download task definitions
sforge fetch-tasks edgebench

# 2. Pull pre-built Docker images
sforge pull --task ad_placement_optimization --registry seededge

# 3. Start judge server (separate terminal)
sforge serve

# 4. Run an agent
SFORGE_AGENT_API_KEY="sk-xxx" \
  sforge run --task ad_placement_optimization --agent claude-code \
    --model "claude-opus-4-8[1m]" --timeout 43200 --run-id edgebench-001

Step-by-step examples:

Important

  • Official setting — leaderboard numbers use the official experiment YAMLs unchanged, including the time budget, stop hook, auto-eval, submission cooldowns, and hardware resource limits.
  • Cost — one 12-hour task with a frontier model can cost hundreds to over a thousand USD; a full ~50-task run is a five-figure spend.
  • Scale — the Docker backend suits only a few tasks at a time; for full-suite runs use the Kubernetes backend.

Evaluating your own model / agent:

  • Your own model — the built-in Claude Code and Codex scaffolds work with any compatible API endpoint: point SFORGE_AGENT_API_BASE_URL at your endpoint, set your key via SFORGE_AGENT_API_KEY, and pass your model name via --model. See Supported Agents.
  • Your own agent scaffold — just add a new agent under sforge/harness/agent/ (a small Agent subclass declaring how to install and launch it) and register it in the factory, then run with --agent <your-agent>. See Custom Agents.

Full documentation: bytedance-seed.github.io/EdgeBench

Citation

If you find EdgeBench useful in your research, please cite our tech report:

@misc{edgebench2026,
  title  = {EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments},
  author = {Deyao Zhu and Xin Zhou and Shengling Qin and Xuekai Zhu and Hangliang Ding and Shu Zhong and others},
  year   = {2026},
  url    = {https://arxiv.org/abs/2607.05155},
}

License

  • EdgeBench Tasks (task datasets) are released under CC BY 4.0.
  • SForge (evaluation harness code) is released under the Apache License 2.0.

Contact

To evaluate on the full 134-task suite, please contact zhongshu@bytedance.com.

ByteDance Seed
Built by ByteDance Seed