Getting Started

August 31, 2026 · View on GitHub

Follow this page start to finish for your first labeling run. You will copy a cookbook, fill in a few machine-specific paths, dry-run it, then run it.

How a run works: a cookbook YAML lists stages, paths, and endpoints. The workflow runner starts one container per stage, in order. Every stage reads and writes the same scene directory under your out_dir. You do not call services yourself unless you are debugging one stage.

If you are changing Auto-Labeling source code, still do this once so you can run the product. Contributor checks live in Local Development.

Before you start

You need:

Step 1 — Confirm the runner is installed

git clone <repo-url>
cd paidf-auto-labeling
make sync
make run SCRIPT=workflow-runner:main ARGS='--help'

If --help fails, stop and fix Installation before continuing.

Step 2 — Provision the models you need

Path A needs GPU checkpoints for detection and (optionally) super resolution. Follow Model Provisioning to download them before continuing — skip this step only if you already have a populated model cache.

Step 3 — Set up a VLM/LLM endpoint

Path A's captioning, Visual QA, and reasoning stages call an external model endpoint. Follow VLM and LLM Endpoints to either deploy one yourself or point at a hosted one, then verify it responds before moving on.

Step 4 — Dry-run the cookbook

A dry run validates cookbook parsing, stage order, image selection, and mount resolution without running any containers. Always do this before a real run — see the Path A commands below.

Step 5 — Run for real

Run the same cookbook without --container-dry-run. This builds/pulls the required images and actually executes each stage.

Step 6 — Confirm the output

A successful run produces one scene directory per input under your configured output root, containing raw/, contextual/, task/, and sidecars/. See Experiment Output Layout for what belongs in each folder.

Once you've completed all six steps for Path A, you know the full shape every other workflow in this repo follows — only the cookbook file and the stages it chains change.

Five Paths

PathWhatNotes
A — Video auto-labelingSuper resolution, tracking, captioning, visual QA, and reasoningBest first cookbook for the general video flow
B — Smart Spaces warehouse reasoningTracking, captioning, event QA, and reasoningBest use-case-specific warehouse sample
C — Video visual attribute search + reasoning + exportTracking, captioning, event QA, reasoning, person QA, visual attribute search, training exportBest end-to-end example with a downstream export
D — Image spatial groundingCaptioning, 2D grounding with SAM3, and referring phrasesImage-only
E — Attribute-only or multi-view visual searchVisual attribute search over explicit attributes or one identity image groupBest attribute-search-specific validation

Prerequisites: Python 3.12+, uv, GNU Make, Docker or Podman, and endpoint or GPU prerequisites for the selected workflow. Full matrix: Installation.

Path A is the worked example below. Paths B–E use the same copy, edit, dry-run, run pattern with a different cookbook.

Path A — Video Auto-Labeling

Cookbook: cookbooks/video_data_augmentation/configs/pipeline_video.yaml

The tracked file uses placeholders for media, SAM3 weights, and endpoint model names. Copy it, then edit only the values for this machine:

cp cookbooks/video_data_augmentation/configs/pipeline_video.yaml \
  cookbooks/video_data_augmentation/configs/pipeline_video.local.yaml

Set these fields in pipeline_video.local.yaml:

FieldWhat to put there
data[*].inputs.media_pathYour video, or a staged NGC traffic clip under data/input_media/videos/
data[*].output.out_dirA writable output directory, for example output/auto_labeling/video_data_augmentation
runtime.model_cache_pathThe checkpoint cache from Model Provisioning
container.mountsRead-only mount of the SAM3 weights into the detection image
endpoints.vlm and endpoints.llmURL and model name from VLM and LLM Endpoints

Stage the NGC traffic clips first (see Samples and Cookbooks). Path A then uses data/input_media/videos/traffic_video_analytics/traffic_sample_000.mp4.

Export the API key in the same shell you will use to run (never put the value in the cookbook):

export NVIDIA_API_KEY="<your-key>"
# Local unauthenticated endpoints can use: export NVIDIA_API_KEY="EMPTY"

Dry-run first. This compiles the stage plan without starting models:

CONFIG=cookbooks/video_data_augmentation/configs/pipeline_video.local.yaml
make run SCRIPT=workflow-runner:main \
  ARGS="--cookbook-file ${CONFIG} --container-dry-run"

Check the printed plan for stage order, image names, the model-cache path, SAM3 mounts, and endpoint URLs. Then run for real:

make run SCRIPT=workflow-runner:main \
  ARGS="--cookbook-file ${CONFIG} --container-user auto --container-ensure-images --container-env NVIDIA_API_KEY"

--container-user auto keeps output files writable across stages. --container-ensure-images builds any missing stage images. --container-env NVIDIA_API_KEY passes the key into every stage container without writing the secret into YAML.

Success: one scene directory appears under your out_dir with raw/, contextual/, task/, and sidecars/. See Experiment Output Layout for what belongs in each folder. If the run fails, start at Troubleshooting.

Path B — Smart Spaces Warehouse Reasoning

Use the Smart Spaces cookbook when you want a use-case-specific warehouse event-verification and reasoning workflow:

  • cookbooks/smart_spaces/configs/pipeline_warehouse_event_reasoning.yaml

It runs:

detection_and_tracking -> captioning -> event_verification_visual_qa -> reasoning

Stage the NGC warehouse clips first. Copy the cookbook to *.local.yaml, set the media path, output path, model cache, SAM3 mount, and endpoint URLs, then run:

CONFIG=cookbooks/smart_spaces/configs/pipeline_warehouse_event_reasoning.local.yaml
make run SCRIPT=workflow-runner:main \
  ARGS="--cookbook-file ${CONFIG} --container-dry-run"

Success signal: the per-scene directory contains event-verification Visual QA sidecars plus reasoning task/ outputs.

Path C — Video Visual Attribute Search, Reasoning, And Training Export

This is the most complete visual-attribute cookbook shipped in the repo:

  • cookbooks/visual_attribute_search/configs/pipeline_video_pas_reasoning.yaml

It runs:

detection_and_tracking
  -> captioning
  -> event_verification_visual_qa
  -> reasoning
  -> person_attribute_visual_qa
  -> person_attribute_search
  -> training_export

Copy it to pipeline_video_pas_reasoning.local.yaml. Set the media path, output path, model cache, SAM3 mount, and endpoint URLs the same way as Path A. Stage the NGC warehouse clips first. Path C then uses data/input_media/videos/warehouse_safety/warehouse_safety_0000.mp4.

CONFIG=cookbooks/visual_attribute_search/configs/pipeline_video_pas_reasoning.local.yaml
export NVIDIA_API_KEY="<your-key>"

make run SCRIPT=workflow-runner:main \
  ARGS="--cookbook-file ${CONFIG} --container-dry-run"

make run SCRIPT=workflow-runner:main \
  ARGS="--cookbook-file ${CONFIG} --container-user auto --container-ensure-images --container-env NVIDIA_API_KEY"

Success:

  • the per-scene directory contains contextual/, task/, and PAS sidecars
  • the configured training-export directory contains a tao-vl-reason-v1.0/ dataset

Video EPAS without export uses cookbooks/visual_attribute_search/configs/pipeline_video_epas.yaml with the same run commands.

Path D — Image Spatial Grounding

Cookbook: cookbooks/image_spatial_grounding/configs/pipeline_grounding.yaml

Copy it to pipeline_grounding.local.yaml, then set:

  • image media_path (sample: data/input_media/images/traffic_intersection_frames/ after you stage NGC traffic clips and extract one still per clip)
  • output directory
  • VLM endpoint
  • SAM3 weights mount

Dry-run and run through the workflow runner as in Path A, including --container-env NVIDIA_API_KEY. Real runs need the SAM3-capable grounding image. pipeline_referring.yaml, in the same cookbooks/image_spatial_grounding/configs/ directory, runs referring_expressions only. Run grounding first, then point referring at the same grounding out_dir so it can read contextual/objects.json.

Path E — Visual-Attribute-Only Image Flows

Two visual-attribute image contracts are checked in:

  • cookbooks/visual_attribute_search/configs/pipeline_image_attributes_pas.yaml — PAS over one attribute JSON file. The attribute-only flow never opens media. Change the JSON path in both media_path and --attribute-json, and keep exactly one data entry.
  • cookbooks/visual_attribute_search/configs/pipeline_image_multiview_pas.yaml — one identity, several views. Person images are not in git. Download RSTPReid (Google Drive archive), copy one identity into data/input_media/images/image_attribute_augmentations/, then keep media_path as a single representative file and pass the folder through --image-group-dir. Change both --image-group-dir paths and media_path together.

Copy to *.local.yaml, set the LLM endpoint, then dry-run and run as in Path A.