Usage and tests

September 17, 2026 · View on GitHub

English README · 中文 README

Run commands from the repository root after completing the README setup. CLI image paths are relative to the request file.

Requests and answers

An example request file:

{
  "image": "red-circle.png",
  "state": "Judge only what is visible.",
  "questions": {
    "color": {
      "type": "choice",
      "instructions": "What color is the shape?",
      "criteria": {"red": "Red", "blue": "Blue", "other": "Another color"}
    },
    "round": {"type": "noul", "instructions": "Is the shape round?"},
    "redness": {
      "type": "score",
      "instructions": "How much of the shape is red?",
      "criteria": ["None", "Some", "All"]
    }
  }
}
TypeResult
choiceOne declared option and its candidate probability distribution
noulnoul, the normalized probability of the “yes” option
scoreProbability-weighted level index, from 0 to number of levels - 1

The program assembles the result; the model does not generate JSON. Probabilities are conditional on the supplied candidates, not calibrated correctness estimates. concentration describes distribution concentration, not accuracy.

Use POST /v1/judge with the same schema, replacing image with a base64 data URL such as data:image/png;base64,.... HTTP does not accept local paths or remote image URLs. The browser handles image encoding for you. See the HTTP example for a complete client.

from jev_visual import Request
from jev_visual.engine import Engine

engine = Engine(".models/Qwen3.5-0.8B-4bit")
result = engine.judge(Request(
    image="examples/dog.jpg",
    questions={"dog": {"type": "noul", "instructions": "Is a dog visible?"}},
))
print(result["answers"]["dog"]["noul"])

Create and use Engine on the same thread. The HTTP server uses a dedicated inference thread and queues requests.

Tests and layout

python -m pytest -q                   # No model download or GPU inference
python -m examples.evaluate           # Real-model visual/cache checks
python -m examples.verify_scoring     # Full-forward oracle for candidate scoring
# With the local server running:
python -m examples.http_smoke

Test outputs go to ignored artifacts/. The committed benchmark and scoring verification preserve the published evidence.

jev_visual/       inference engine, adapters, schema, CLI and local server
examples/         runnable requests, image fixtures and integration checks
tests/            unit and API contract tests
benchmarks/       runners, methodology and measured results
docs/             candidate scoring and implementation details
third_party/      upstream license notices