AgentBench Evaluation Results

July 24, 2026 · View on GitHub

Chinese

Agent Memory Evaluation guide · Architecture

These tables record OpenClaw and Hermes evaluation setups. The current public Agent Memory runner guide and registered runtime adapter cover OpenClaw.

Metrics

Acc: the average single-run pass rate across 3 independent runs of the same task, used to measure stable task completion in one run.

Avg turns: the average number of model response turns triggered by the agent per task, used to measure interaction and reasoning depth.

Avg chars: the average character length of the agent's response text, used to measure output length.

Data And Evaluation Setup

Dataset source and split logic:

Reference: https://huggingface.co/datasets/EverMind-AI/EvoAgentBench

Evaluated agents: OpenClaw and Hermes with the corresponding product plugin or integration.

Baseline refers to the corresponding evaluated agent running without any plugin.

The OpenClaw version used for OpenClaw evaluation is 2026.5.7.

The Hermes version used for Hermes evaluation is Hermes Agent v0.18.0 (2026.7.1).

Products not marked as cloud services use locally deployed services and their corresponding agent plugins or integrations.

MemOS uses Memos-Local-Plugin version 2.0.8.

OpenClaw and Hermes answer model configuration: qwen3.6-flash in no_thinking mode.

Evaluation judge model: qwen3.6-flash in thinking mode.

Results

OpenClaw

MethodBrowseComp-Plus AccBrowseComp-Plus Avg turnsOmniMath
Acc
OmniMath Cost
Avg chars
SWE-Bench
Acc
SWE-Bench
Avg turns
LiveCodeBench
Acc
LiveCodeBench Avg turnsGDPVal
Acc
GDPVal
Avg turns
Baseline18.4635.1525658.626.9258.851.2823.934.4817.2
Mem013.3336.4536090.5326.9255.7860.686.2741.3817.28
EverOS18.9830.854.676085.8732.0554.7759.837.238.5116.97
Supermemory14.3632.03585632.130.751.3759.836.8343.1016.70
Hindsight16.9235.655617738.4656.771.7914.250.0017.5
mem9
(Cloud service)
12.3131.654.335880.930.7759.164.18.2347.715.8
Memori
(Cloud service)
16.9236.3526189.234.6243.151.289.337.9315.8
MemOS23.85*43.361*5164.238.46*49.5564.969.4362.07*23.93

Hermes

MethodBrowseComp-Plus AccBrowseComp-Plus Avg turnsOmniMath
Acc
OmniMath Cost
Avg chars
SWE-Bench
Acc
SWE-Bench
Avg turns
LiveCodeBench
Acc
LiveCodeBench Avg turnsGDPVal
Acc
GDPVal
Avg turns
Baseline18.9743.236210872.837.1884.9352.9922.2754.662.57
Mem017.9542.7364.331174034.6294.3757.2625.6356.957.73
Supermemory18.4641.0364.0010675.9035.9092.2366.6719.7357.4758.60
Hindsight21.5448.3364.0012089.6328.2084.4761.5430.0757.4759.67
Viking21.5439.3358.6712514.2035.9088.7060.6827.7758.0554.27
mem9
(Cloud service)
18.4640.136612419.439.7498.458.9718.7755.7553.17
Memori
(Cloud service)
14.8740.1364.331481128.2191.554.722.3755.7553.17
MemOS20.7536.2072.67*5915.752.56*76.3764.125.6355.1758.6