Evaluation of Pose Estimation Model
January 22, 2026 ยท View on GitHub
Our evaluation service is a comprehensive tool that enables users to assess the accuracy of their TensorFlow Lite ( .tflite) or Keras (.keras) pose estimation model. By uploading their model and a validation set, users can quickly and easily evaluate the performance of their model and generate various metrics, such as OKS or mAP.
The evaluation service is designed to be fast, efficient, and accurate, making it an essential tool for anyone looking to evaluate the performance of their pose estimation model.
1. Configure the YAML file
To use this service and achieve your goals, you can use the user_config.yaml or directly update the evaluation_config.yaml file and use it. This file provides an example of how to configure the evaluation service to meet your specific needs.
Alternatively, you can follow the tutorial below, which shows how to evaluate your pre-trained pose estimation model using our evaluation service.
.keras: Tensorflow/keras float models that can be used in prediction, training, evaluation, and quantization services..tflite: Tensorflow/keras quantized models that can be used in prediction, evaluation, benchmark, and deployment services..onnx(float): Open Neural Network Exchange float models that can be used in prediction, evaluation, and quantization services..onnx(qdq): Open Neural Network Exchange quantized models that can be used in prediction, evaluation, benchmark, and deployment services.-
spe: These are single pose estimation models that output directly the keypoints positions and confidences. -
hand_spe: These are single hand landmarks estimation models that outputs directly the keypoints positions and confidences of the hand pose. -
head_spe: These are single head landmarks estimation models that outputs directly the keypoints positions and confidences of the head pose. -
heatmaps_spe: These are single pose estimation models that outputs heatmaps that we must post-process in order to get the keypoints positions and confidences. -
yolo_mpe: These are the YOLO (You Only Look Once) multiple pose estimation models from Ultralytics that outputs the same tensor as in object detection but with the addition of a set of keypoints for each bbox.
1.1 Set the model and the operation mode
As mentioned previously, all the sections of the YAML file must be set in accordance with
this YAML file.
In particular, operation_mode should be set to evaluation and the evaluation section should be filled as in the
following example:
model:
model_path: https://github.com/stm32-hotspot/ultralytics/raw/refs/heads/main/examples/YOLOv8-STEdgeAI/stedgeai_models/pose_estimation/yolov8n_256_quant_pt_uf_pose_coco-st.tflite
model_type: yolo_mpe
operation_mode: evaluation
In this example, the path to the yolov8n multi pose estimation model is provided in the model_path parameter.
But you can provide any of these types :
The model_type attribute specifies the type of the model architecture that you want to train. It is important to note
that only certain models are supported. These models include:
evaluation:
target: host # host, stedgeai_host, stedgeai_n6
In the 'evaluation' section, if users are using a quantized TFLITE or ONNX model, they can decide to do the inferences with the classic python interpreters (host -> by default), with the C code generated by stedgeai on the PC (stedgeai_host), or with the C code generated by stedgeai on the N6 board directly (stedgeai_n6) using the target attribute.
1.2 Prepare the dataset
Information about the dataset you want use for evaluation is provided in the dataset section of the configuration
file, as shown in the YAML code below.
dataset:
dataset_name: coco # Dataset name/type
keypoints: 17 # number of keypoints of each pose
test_path: <test-set-root-directory> # Path to the root directory of the test set.
In this example, the path to the validation set is provided in the test_path parameter.
State machine below describes the rules to follow when handling dataset path for the evaluation.
Important
In 'dataset' section, the dataset_name is mandatory and should always be set to 'coco' even if you dont use the COCO dataset
In cases where there is no validation set path or test set provided to evaluate the model trained using the training service, the available data under the training_path directory is split into two to create a training set and a validation set. By default, 80% of the data is used for training and the remaining 20% is used for the validation set in the evaluation service.
If you want to use a different split ratio, you need to specify the percentage to be used for the validation set in the validation_split parameter, you must specify the same validation_split parameter value in both the training and evaluation services, as shown in the YAML example below:
dataset:
dataset_name: coco
training_path: ../datasets/COCO_2017_pose/
validation_path:
validation_split: 0.20
test_path:
- The value of
interpolationmust be one of {"bilinear", "nearest", "bicubic", "area", "lanczos3", "lanczos5", "gaussian", "mitchellcubic"}. - The value of
aspect_ratiomust have a value of "fit" as we do not support other values such as "crop". If you set it to "fit", the resized images will be distorted if their original aspect ratio is not the same as in the resizing size.
1.3 Apply preprocessing
The images from the dataset need to be preprocessed before they are presented to the network for evaluation. This includes rescaling and resizing. In particular, they need to be rescaled exactly as they were at training step. This is illustrated in the YAML code below:
preprocessing:
rescaling: { scale: 1/127.5, offset: -1 }
resizing:
aspect_ratio: "fit"
interpolation: nearest
color_mode: rgb
In this example, the pixels of the input images read in the dataset are in the interval [0, 255], that is UINT8. If you
set scale to 1./255 and offset to 0, they will be rescaled to the interval [0.0, 1.0].
If you set scale to 1/127.5 and offset to -1, they will be rescaled to the interval [-1.0, 1.0].
The resizing attribute specifies the image resizing methods you want to use:
The color_mode attribute must be one of "grayscale", "rgb" or "rgba".
When you define the preprocessing parameter in the configuration file, the annotation file for the pose estimation
dataset will be automatically modified during preprocessing to ensure that it is aligned with the preprocessed images.
This typically involves updating the bounding box coordinates to reflect any resizing or cropping that was performed
during preprocessing.
This automatic modification of the annotations file is an important step in preparing
the dataset for pose estimation, as it ensures that the annotations accurately reflect the preprocessed images and
enables the model to learn from the annotated data.
confidence_thresh- A float between 0.0 and 1.0, the score threshold to filter detections.NMS_thresh- A float between 0.0 and 1.0, NMS threshold to filter and reduce overlapped boxes.max_detection_boxes- An int between 0 and infinity, the maximum number of poses that the multi-pose can output for one image.plot_metrics- A bool that allows to plot the mAP curves
1.4 Apply post-processing
postprocessing:
confidence_thresh: 0.001
NMS_thresh: 0.1
max_detection_boxes: 100
plot_metrics: true
Apply post-processing by modifiying the postprocessing parameters in user_config.yaml as the following:
1.5 Hydra and MLflow settings
The mlflow and hydra sections must always be present in the YAML configuration file. The hydra section can be used
to specify the name of the directory where experiment directories are saved and/or the pattern used to name experiment
directories. With the YAML code below, every time you run the Model Zoo, an experiment directory is created that
contains all the directories and files created during the run. The names of experiment directories are all unique as
they are based on the date and time of the run.
hydra:
run:
dir: ./src/experiments_outputs/${now:%Y_%m_%d_%H_%M_%S}
The mlflow section is used to specify the location and name of the directory where MLflow files are saved, as shown
below:
mlflow:
uri: ./src/experiments_outputs/mlruns
2. Evaluate your model
If you chose to modify the user_config.yaml you can evaluate the model by running the following command from the UC folder:
python stm32ai_main.py
If you chose to update the evaluation_config.yaml and use it then run the following command from the UC folder:
python stm32ai_main.py --config-path ./config_file_examples/ --config-name evaluation_config.yaml
In case you want to evaluate the accuracy of the quantized model then benchmark it, you can either launch the evaluation operation mode followed by the benchmark service that describes in detail how to proceed or you can use chained services like launching chain_eqeb example with the command below:
python stm32ai_main.py --config-path ./config_file_examples/ --config-name chain_eqeb_config.yaml
3. Visualize the evaluation results
You can retrieve the confusion matrix generated after evaluating the float/quantized model on the test set by navigating to the appropriate directory within experiments_outputs/<date-and-time>.
You can also find the evaluation results saved in the log file stm32ai_main.log under ** experiments_outputs/<date-and-time>**.