Claw-Eval Adapter
August 19, 2026 · View on GitHub
This directory contains the OpenClaw execution adapter for the TokenPilot runtime path on the claw-eval benchmark.
The harness invokes the local LightRSI release installer through TOKENPILOT_RUNTIME_ROOT, while OpenClaw loads the installed runtime copy from ~/.openclaw/extensions/tokenpilot.
The current layout is designed to be mostly self-contained inside this repo:
scripts/: runtime adapter entrypoints and execution gluedataset/tasks/: local task source of truthdataset/general/: flatgeneralasset bundle locationvendor/claw_eval_src/: vendored upstreamclaw_evalPython packagevendor/mock_services/: vendored upstream mock servicesplugins/: vendoredclaw-eval-mock-tools*OpenClaw plugins
What is already vendored
The adapter no longer requires the external upstream repo checkout at runtime for core execution:
- upstream
src/claw_evalis vendored undervendor/claw_eval_src/ - upstream
mock_servicesis vendored undervendor/mock_services/ - benchmark mock plugins are vendored under
plugins/
Remaining external requirements
A fresh clone is not enough by itself. You still need a working OpenClaw runtime environment.
Required environment pieces:
- OpenClaw installed and usable from the shell
- a valid OpenClaw home/config, typically under
~/.openclaw/ - provider API keys / model routes configured in
openclaw.json - the flattened
generalasset bundle copied intodataset/general/
External data layout
Large Claw-Eval assets are stored outside Git.
Google Drive root:
Recommended Drive layout for this benchmark:
LightRSI-Experiment/
└── TokenPilot/
├── experiment-data/
│ └── claw-eval/
│ ├── general/
│ └── tasks/
└── experiment-results/
└── claw-eval/
Local mount points in this repository:
benchmarks/claw-eval/dataset/general/benchmarks/claw-eval/dataset/tasks/
See:
When you prepare a fresh machine:
- download
general/from Drive - copy its contents into
dataset/general/ - download
tasks/from Drive - copy it back under
dataset/tasks/
Official runners
The recommended public runner surface is intentionally small:
scripts/run_baseline.shscripts/run_method.sh
These are the main entrypoints for reproduction.
Baseline
Minimal isolated baseline smoke:
cd /path/to/TokenPilot
export TOKENPILOT_RUNTIME_ROOT=/path/to/LightRSI
bash benchmarks/claw-eval/scripts/run_baseline.sh \
--scope suite \
--suite T001zh_email_triage \
--session-mode isolated \
--model gpt-5.4-mini
Run all general tasks in isolated baseline mode:
cd /path/to/TokenPilot
export TOKENPILOT_RUNTIME_ROOT=/path/to/LightRSI
bash benchmarks/claw-eval/scripts/run_baseline.sh \
--scope general \
--session-mode isolated \
--model gpt-5.4-mini
Method
Minimal isolated method smoke:
cd /path/to/TokenPilot
export TOKENPILOT_RUNTIME_ROOT=/path/to/LightRSI
bash benchmarks/claw-eval/scripts/run_method.sh \
--scope suite \
--suite T001zh_email_triage \
--session-mode isolated \
--profile reduction \
--model lightrsi/gpt-5.4-mini
Run all general categories in continuous method mode:
cd /path/to/TokenPilot
export TOKENPILOT_RUNTIME_ROOT=/path/to/LightRSI
bash benchmarks/claw-eval/scripts/run_method.sh \
--scope general \
--session-mode continuous \
--profile plugin \
--by-category \
--model lightrsi/gpt-5.4-mini
If your primary OpenClaw config is read-only or you want run-local isolation, add --tmp-openclaw:
cd /path/to/TokenPilot
export TOKENPILOT_RUNTIME_ROOT=/path/to/LightRSI
bash benchmarks/claw-eval/scripts/run_method.sh \
--scope general \
--session-mode continuous \
--profile plugin \
--by-category \
--tmp-openclaw \
--model lightrsi/gpt-5.4-mini
Current state
What is in good shape:
- task loading and suite selection
- isolated execution path
- continuous execution path
- upstream grader bridge
- vendored upstream runtime code
- vendored mock service plugins
- repo-internal default paths for code, mock services, plugins, and
generalassets - unified baseline/method runner surface
What is still operationally sensitive:
- OpenClaw plugin/config state in
~/.openclaw/openclaw.json - provider stability / request timeouts
- duplicate plugin ids from previously installed local extensions
- plugin continuous experiments with the current TokenPilot component reduction/eviction/estimator enabled
- large fixture synchronization between local working copies and the shared Drive mirror
Known pitfalls
1. dataset/general/ must be populated
If the flat general bundle is missing, file-backed tasks will fail.
2. OpenClaw config pollution
Previous runs can leave stale plugin allowlists or entries in ~/.openclaw/openclaw.json. This can break both claw-eval and pinchbench runs. If you want a safer run-local copy, use --tmp-openclaw on the official runners.
3. Duplicate plugin ids
If the same plugin exists both in the vendored plugins/ directory and in a previously installed OpenClaw extension path, OpenClaw may warn about duplicate plugin ids.
4. Provider/runtime stability is separate from repository layout
A repo-internal smoke can still fail because of provider timeouts or runtime environment issues even when all local paths are correct.