Hermes Backend Benchmark
March 29, 2026 ยท View on GitHub
Status:
implemented_earlyon 2026-03-28 for same-host Hermes benchmark evidence against Psionic, Ollama, and an optionalllama.cppcomparator lane.
This document records the retained same-host Hermes backend benchmark for local
qwen35, keeping Hermes fixed and swapping only the OpenAI-compatible backend
endpoint and backend-specific model identifier.
The repo now retains three benchmark receipts:
- the first Psionic-vs-Ollama benchmark at
473c0ad5bd5219cbcb7f76495d8166933242b872 - the optional three-way follow-on at
477ca8226589a5b1760ed93c47b87041f08172ab - the serialized two-city extension at
20260329_archlinux_2b_serialized
The later receipt is the canonical answer for the optional llama.cpp
comparator issue because it preserves the same benchmark contract and records
the third lane honestly even when that lane is unavailable on the current host.
The latest serialized receipt is the canonical answer for the consumer-GPU two-city Hermes path because it adds the overlooked serialized workflow and retains the Psionic-vs-Ollama split honestly.
Exact Revisions
- Psionic revision:
e1a27665c63f4b7a085795415699ab7f168c4d2d - Hermes revision:
e295a2215acd55f2ee930fc7a4cd2df1c5464234 - Host:
archlinux llama-serverversion:7444 (58062860a)
Canonical Runner
Run the benchmark from the repo root:
scripts/release/run-hermes-backend-benchmark.sh
The exact retained three-way receipt used:
- Psionic model path:
/home/christopherdavid/models/qwen3.5/qwen3.5-2b-q8_0-registry.gguf - Ollama model:
qwen3.5:2b - Hermes root:
/home/christopherdavid/scratch/hermes-agent-proof2 - Hermes Python:
/home/christopherdavid/scratch/hermes-min/.venv/bin/python - Psionic server binary:
/home/christopherdavid/.cache/psionic-hermes-target/debug/psionic-openai-server llama.cppserver binary:/home/christopherdavid/code/llama.cpp/build/bin/llama-server
Exact retained command:
TMPDIR=/home/christopherdavid/scratch/tmp/hermes-bench \
PSIONIC_HERMES_ROOT=/home/christopherdavid/scratch/hermes-agent-proof2 \
PSIONIC_HERMES_PYTHON=/home/christopherdavid/scratch/hermes-min/.venv/bin/python \
PSIONIC_HERMES_SERVER_BIN=/home/christopherdavid/.cache/psionic-hermes-target/debug/psionic-openai-server \
PSIONIC_HERMES_PSIONIC_MODEL_PATH=/home/christopherdavid/models/qwen3.5/qwen3.5-2b-q8_0-registry.gguf \
PSIONIC_HERMES_OLLAMA_MODEL=qwen3.5:2b \
PSIONIC_HERMES_ENABLE_LLAMACPP=1 \
PSIONIC_HERMES_LLAMA_CPP_MODEL_ALIAS=qwen3.5-2b-q8_0-registry.gguf \
PSIONIC_HERMES_BENCHMARK_REPORT_PATH=/home/christopherdavid/scratch/psionic-hermes-benchmark-477ca822/fixtures/qwen35/hermes/hermes_psionic_vs_ollama_vs_llamacpp_benchmark_20260328_archlinux_2b.json \
PSIONIC_HERMES_BENCHMARK_RAW_DIR=/home/christopherdavid/scratch/psionic-hermes-benchmark-477ca822/fixtures/qwen35/hermes/backend_rows \
scripts/release/run-hermes-backend-benchmark.sh
The exact retained serialized rerun used:
TMPDIR=/home/christopherdavid/scratch/tmp/hermes-bench-serialized \
PSIONIC_HERMES_SKIP_BUILD=1 \
PSIONIC_HERMES_ROOT=/home/christopherdavid/scratch/hermes-agent-proof2 \
PSIONIC_HERMES_PYTHON=/home/christopherdavid/scratch/hermes-min/.venv/bin/python \
PSIONIC_HERMES_SERVER_BIN=/home/christopherdavid/.cache/psionic-hermes-target/debug/psionic-openai-server \
PSIONIC_HERMES_PSIONIC_MODEL_PATH=/home/christopherdavid/models/qwen3.5/qwen3.5-2b-q8_0-registry.gguf \
PSIONIC_HERMES_OLLAMA_MODEL=qwen3.5:2b \
PSIONIC_HERMES_REQUEST_TIMEOUT_SECONDS=60 \
PSIONIC_HERMES_BENCHMARK_REPORT_PATH=/home/christopherdavid/scratch/psionic-hermes-664/fixtures/qwen35/hermes/hermes_psionic_vs_ollama_benchmark_20260329_archlinux_2b_serialized.json \
PSIONIC_HERMES_BENCHMARK_RAW_DIR=/home/christopherdavid/scratch/psionic-hermes-664/fixtures/qwen35/hermes/backend_rows \
PSIONIC_HERMES_SERVER_LOG_PATH=/home/christopherdavid/scratch/logs/hermes_backend_benchmark_20260329_archlinux_2b_serialized.log \
scripts/release/run-hermes-backend-benchmark.sh
Fixed Contract
This benchmark keeps the following fixed:
- same host
- same Hermes revision
- same Hermes
chat.completionscustom-provider path - same case ids
- same tool schemas and handlers
temperature = 0seed = 0- same
required_then_autotool policy for tool cases - same
autopolicy for the no-tool case
What changed between rows:
OPENAI_BASE_URL- backend-specific model identifier
- Psionic: local GGUF basename
- Ollama: local Ollama model name
llama.cpp: local GGUF basename
Retained Reports
- first two-way aggregate:
fixtures/qwen35/hermes/hermes_psionic_vs_ollama_benchmark_20260328_archlinux_2b.json - three-way aggregate:
fixtures/qwen35/hermes/hermes_psionic_vs_ollama_vs_llamacpp_benchmark_20260328_archlinux_2b.json - serialized two-way aggregate:
fixtures/qwen35/hermes/hermes_psionic_vs_ollama_benchmark_20260329_archlinux_2b_serialized.json - three-way Psionic row:
fixtures/qwen35/hermes/backend_rows/hermes_psionic_row_20260328_archlinux_2b_llamacpp_comparator.json - three-way Ollama row:
fixtures/qwen35/hermes/backend_rows/hermes_ollama_row_20260328_archlinux_2b_llamacpp_comparator.json - serialized Psionic row:
fixtures/qwen35/hermes/backend_rows/hermes_psionic_row_20260329_archlinux_2b_serialized.json - serialized Ollama row:
fixtures/qwen35/hermes/backend_rows/hermes_ollama_row_20260329_archlinux_2b_serialized.json - three-way
llama.cpprow:fixtures/qwen35/hermes/backend_rows/hermes_llama_cpp_row_20260328_archlinux_2b.json
Cases
The retained benchmark now uses five Hermes cases:
auto_plain_text_turnrequired_tool_turnmulti_turn_tool_loopstreamed_tool_turnserialized_two_city_tool_loop
Psionic and Ollama both passed the original four retained cases on the exact
477ca822 rerun. The later serialized receipt extends that same-host contract
with the overlooked consumer-GPU two-city lane.
Serialized Two-City Extension
The serialized receipt adds the consumer-GPU workflow that matters for Hermes operability on one local GPU:
- Hermes must call
get_paris_weather - then Hermes must call
get_tokyo_weather - then the harness asks the same backend for the exact final sentence
- the direct summary closeout is bounded with
PSIONIC_HERMES_REQUEST_TIMEOUT_SECONDS=60
This bound matters because the earlier benchmark extension could hang forever on the Ollama side and fail to emit any aggregate report. The repo-owned probe now records that timeout as a retained failure detail instead of stalling the whole benchmark.
Serialized result on archlinux RTX 4080:
| Case | Psionic | Ollama | Honest Result |
|---|---|---|---|
serialized_two_city_tool_loop | pass, 5.0134s | fail, 63.2206s | Psionic completes the full serialized path; Ollama completes the tool loop but times out in the bounded direct summary phase |
Per-row serialized summary:
- Psionic:
overall_pass = truepassing_case_count = 5/5mean_case_wallclock_s = 2.9890- serialized direct-summary
completion_tok_s = 39.4327
- Ollama:
overall_pass = falsepassing_case_count = 4/5mean_case_wallclock_s = 13.9439- serialized failure detail:
summary_phase.error_type = TimeoutErrorsummary_phase.error_message = timed outafter60.0631s
That retained difference is the point of the issue: keeping Hermes fixed and swapping only the backend now makes the serialized consumer-GPU gap explicit.
Four-Case Baseline Result
| Case | Psionic wallclock s | Ollama wallclock s | llama.cpp | Faster |
|---|---|---|---|---|
required_tool_turn | 3.1804 | 2.9847 | startup_failure | ollama |
auto_plain_text_turn | 1.1518 | 0.4791 | startup_failure | ollama |
multi_turn_tool_loop | 3.1403 | 1.7067 | startup_failure | ollama |
streamed_tool_turn | 2.9973 | 1.7697 | startup_failure | ollama |
Per-row baseline summary:
- Psionic:
overall_pass = truepassing_case_count = 4/4mean_case_wallclock_s = 2.6174mean_completion_tok_s = null
- Ollama:
overall_pass = truepassing_case_count = 4/4mean_case_wallclock_s = 1.7351mean_completion_tok_s = 88.5821
llama.cpp:overall_pass = falserow_status = startup_failurefailure_detail = llama.cpp server failed readiness check on 127.0.0.1:8098
Baseline availability probe on this same host:
- Psionic readiness probe:
0.8806s - Ollama readiness probe:
0.0127s llama.cppreadiness probe:nullllama.cppstartup status:startup_failure
That readiness probe is not a cold-start apples-to-apples claim. Psionic is being launched by the harness, while Ollama is measured as an already-running same-host service.
Why The llama.cpp Lane Matters
This lane matters only when llama.cpp is a real runnable backend for the same
model contract. That is useful for:
- Mac-local CPU or Metal operability questions
- checking whether the same model artifact can cross the Psionic and
llama.cppboundary without format or tokenizer drift - separating "Psionic is slower" from "the comparator is not actually runnable here"
On this exact archlinux RTX 4080 host, the lane is currently a compatibility
boundary check, not a throughput comparator:
- the local
llama-serverbinary can start, but it cannot load the retained rewritten Psionicqwen35GGUF - the retained
llama.cpprow captures the exact failure:unknown model architecture: 'qwen35' ollama show --modelfile qwen3.5:2bpoints at/usr/share/ollama/.ollama/models/blobs/sha256-b709d81508a078a686961de6ca07a953b895d9b286c46e17f00fb267f4f2d297- that blob path is permission-blocked to the current user, so the harness
cannot honestly reroute the Ollama-managed artifact through
llama.cppon this host
That is why the wrapper now retains a synthetic llama_cpp row instead of
crashing the whole benchmark. The comparator boundary is now explicit,
machine-readable, and rerunnable.
Honest Interpretation
The retained three-way receipt does not show usable llama.cpp throughput
for this qwen3.5 lane. It shows:
- Psionic and Ollama both complete the same four Hermes cases on the same host under the same high-level contract
- Ollama still wins wallclock on all four retained cases
- the optional
llama.cppcomparator is wired into the same harness, but the current host and model contract stop it at startup
That is enough to close the optional comparator issue honestly. The repo no longer has a silent third-lane gap or a benchmark wrapper that aborts as soon as the comparator cannot start.
The newer serialized receipt adds a second important conclusion:
- Psionic is still slower than Ollama on the four older easy same-host Hermes cases
- but Psionic currently wins the harder consumer-GPU serialized two-city lane because it completes the bounded direct-summary closeout while Ollama times out there
So the current backend truth is not "Ollama simply wins everything locally." The retained consumer-GPU serialized path now shows a real task where Psionic has the better backend behavior under the same Hermes controller.
Current Gaps
Two real gaps remain visible after the three-way receipt:
- Psionic is still slower than Ollama on all four current same-host Hermes
cases on this
qwen3.52brow - Psionic still does not expose comparable usage accounting on this local
Hermes benchmark lane, so
completion_tok_sremainsnull
The optional llama.cpp lane also remains blocked for real throughput
comparison until one of these becomes true:
llama.cppgains support for the rewrittenqwen35GGUF contract- the benchmark can access a
llama.cpp-loadable model artifact on the same host without permission tricks or hidden format swaps
Relation To Other Hermes Issues
The repeated same-backend warm-path reuse proof now lives in
docs/HERMES_QWEN35_REUSE_BENCHMARK.md.
- direct compatibility proof remains in
docs/HERMES_QWEN35_COMPATIBILITY.md - fast-path-versus-fallback runtime truth now lives in
docs/HERMES_QWEN35_FAST_PATH.md - session reuse and prefix-cache reuse belong to the later repeated-session latency issue