Contributing to the Reference Scenarios
August 27, 2026 · View on GitHub
This directory contains the runnable reference scenarios and the tooling used to validate them against the GenAI semantic conventions.
If you are changing the semantic conventions themselves under model/ or
docs/, use the repository-level guide in ../CONTRIBUTING.md.
Structure
pyproject.toml # Tooling project metadata
src/
semconv_genai/ # Scenario discovery, weaver install, report generation
scenarios/
<library>/ # Reference scenarios
Within each scenario directory:
conformance.yaml— How to run the scenario and what it must producescenario.py— SDK invocation + manual OTel spanspyproject.toml— Dependenciesuv.lock— Locked transitive dependency graph (committed)data.json— Committed results
The conformance runner
Scenarios run under the
conformance runner,
which owns the mock LLM server, the weaver live-check lifecycle, the GenAI
advice policies, and the reduction that produces data.json. None of it is
vendored here: it is fetched at the CONFORMANCE_REF pinned in the
repository-root versions.env and cached under
~/.cache/semconv-genai/conformance/.
Runs are checked against the working tree's model/, not the registry that
package pins, so a change to the conventions takes effect immediately.
Violations are reported as warnings, not failures — the scenarios are not yet
clean against the conventions and their gaps are not yet declared under
expected_violations. Pass --strict to fail on them. A scenario that crashes
or misses what its conformance.yaml declares fails either way.
Prerequisites
- uv (uv will fetch the Python 3.12 interpreter declared in
pyproject.tomlon first run).
Run the commands below from this reference/ directory.
Running scenarios
First-time setup creates .venv and installs the tooling:
uv sync
Run a single library, or all libraries:
uv run run-scenario openai # one library
uv run run-scenario --all # all libraries
uv run run-scenario --all --keep-going # continue through failures, report at end
uv run run-scenario <library> runs the selected scenario under
scenarios/ against a local mock LLM server, validates the
emitted telemetry, and writes the results that feed the checked-in reports.
The raw weaver report for each run is left under scenarios/<library>/output/.
Linting
Lint and format the Python code under src/semconv_genai/ and scenarios/:
uv tool run --from ruff ruff check --fix src/semconv_genai scenarios
uv tool run --from ruff ruff format src/semconv_genai scenarios
Updating reports
Regenerate the checked-in status section in README.md after updating committed
data.json files:
uv run update-reports
Contribution expectations
- Keep reference coverage honest. Only emit signals and attributes that the library or reference code can actually produce.
- Follow the authoring guidelines and patterns in the reference skill.
- Prefer focused updates to the affected library under
scenarios/<library>/. - After regenerating
scenarios/*/data.json, runuv run update-reportsand commit both alongside your change.
Which operations a scenario should emit
A scenario emits telemetry only for the operations the library itself performs. Two principles:
- Don't re-emit another library's telemetry. If the library delegates an
operation to another instrumentable library (for example a framework that
calls
openai,anthropic, orgoogle-genaiunder the hood), instrumentation for that operation belongs to the underlying library. - Emit an operation only when the library has that concept. Emit inference or embeddings only when the library is itself the model-call boundary, an agent span only when it models agents, a workflow span only when it models workflows or graphs, and so on. Calling the provider's REST API directly (with no separate instrumentable client library in between) makes the library the model-call boundary, so it is a valid reason to emit inference.
Agent-framework instrumentation SHOULD NOT emit inference spans by default; the underlying LLM library owns them. The exception is a framework that issues the model call in a way no other instrumentable library can observe — it calls the REST API directly, embeds a vendored model library, or similar. In that case the framework is the only place the inference call is visible, so it SHOULD emit the span.
If a library emits unrelated native telemetry that obscures the intended validation surface, suppress that library-owned telemetry in the reference scenario rather than changing the semantic conventions to match it.
Adding or updating a library
Reference scenarios are both validation inputs and examples for instrumentation authors, so keep them minimal and readable.
When adding a new reference scenario:
-
Create
scenarios/<library>/scenario.py. -
Copy
conformance.yamlfrom a neighbouring scenario and setinstrumented_libraryto the library slug. -
Create
scenarios/<library>/pyproject.tomldeclaring the SDK dependencies plusgenai-reference-shared(sourced from the shared project atshared/). The OTel SDK pin is provided transitively bygenai-reference-shared; do not re-declare it here unless the library needs a non-default version.[project] name = "<library>-reference-test" version = "0" requires-python = ">=3.12" dependencies = [ "<sdk>==<pinned-version>", "genai-reference-shared", ] [tool.uv.sources] genai-reference-shared = { path = "../../shared", editable = true } [tool.uv] package = falseIf the SDK under test requires a specific OTel version that differs from the default pin in shared/pyproject.toml, add
override-dependenciesto the existing[tool.uv]table:[tool.uv] package = false override-dependencies = [ "opentelemetry-api==<required>", "opentelemetry-sdk==<required>", "opentelemetry-exporter-otlp-proto-grpc==<required>", ] -
Run
uv lockinsidescenarios/<library>/to generate the committeduv.lock. Re-run it whenever you change dependencies; the scenario'sruncommand usesuv run --frozenand will fail if the lockfile is stale. -
Run
uv run run-scenario <library>to generatescenarios/<library>/data.json. -
Regenerate
README.mdwithuv run update-reports.