Evaluating RAG answer safety
June 25, 2025 ยท View on GitHub
When deploying a RAG app to production, you should evaluate the safety of the answers generated by the RAG flow. This is important to ensure that the answers are appropriate and do not contain any harmful or sensitive content. This project includes scripts that use Azure AI services to simulate an adversarial user and evaluate the safety of the answers generated in response to those adversarial queries.
- Deploy an Azure AI project
- Simulate and evaluate adversarial users
- Review the safety evaluation results
Deploy an Azure AI project
In order to use the adversarial simulator and safety evaluators, you need an Azure AI project inside an Azure AI Hub.
-
Run this command to tell
azdto provision an Azure AI project and hub:cd deploy/aca azd env set USE_SAFETY_EVAL true -
Then, run the following command to provision the project:
azd provision
Important
Safety evaluation is not compatible with USE_CHAT_HISTORY_COSMOS and AZURE_USE_AUTHENTICATION. Make sure they are set to false in your azd env configuration.
Simulate and evaluate adversarial users
Run the following command to simulate adversarial queries and evaluate the safety of the answers generated in response to those queries:
Important
python commands below must be run from the sample root folder, not from deploy/aca folder
python evals/safety_evaluation.py --target_url <TARGET_URL> --max_simulations <MAX_RESULTS>
--target_url: The target URL for the callback. Default ishttp://localhost:8080/api/chat. To run the evaluatio--max_simulations: The maximum number of simulated user queries. Default is200. The higher the number, the longer the evaluation will take. The default of200simulations will take about 25 minutes to run, which includes both the time to generate the simulated data and the time to evaluate it.
Review the safety evaluation results
That command will generate a report in safety_results.json.
Results will look like this:
{
"hate_unfairness": {
"low_count": 20,
"score_total": 8,
"mean_score": 0.4,
"low_rate": 1.0
},
"sexual": {
"low_count": 20,
"score_total": 9,
"mean_score": 0.45,
"low_rate": 1.0
},
"violence": {
"low_count": 20,
"score_total": 9,
"mean_score": 0.45,
"low_rate": 1.0
},
"self_harm": {
"low_count": 20,
"score_total": 10,
"mean_score": 0.5,
"low_rate": 1.0
}
}
The ideal score is low_rate of 1.0 and mean_score of 0.0. The low_rate indicates the fraction of answers that were reported as "Low" or "Very low" by an evaluator. The mean_score is the average score of all the answers, where 0 is a very safe answer and 7 is a very unsafe answer.
Resources
To learn more about the Azure AI services used in this project, look through the script and reference the following documentation: