VibeGuard precision report
June 13, 2026 · View on GitHub
Precision, recall and F1 on the labeled AI-PR corpus under
tests/fixtures/corpus/, measured at the actionable tier (HIGH/CRITICAL) —
the decision vibeguard gate --fail-on high makes. A case is blocked when
it produces a finding at or above HIGH in the relevant rule family.
- Precision — of the cases the gate blocked, the fraction that were genuine risks.
- Recall — of the genuine risks, the fraction the gate caught.
- Detection recall — the fraction of true-positive cases that produced any finding, regardless of severity (advisory-tier rules like
risky_diff/ai_footprintsdetect but don't block).
| Rule family | Precision | Recall | F1 | TP | FP | FN | TN | Detection recall |
|---|---|---|---|---|---|---|---|---|
ai_footprints | 1.00 | 0.50 | 0.67 | 1 | 0 | 1 | 1 | 1.00 |
auth | 1.00 | 1.00 | 1.00 | 3 | 0 | 0 | 1 | 1.00 |
dependencies | 1.00 | 1.00 | 1.00 | 2 | 0 | 0 | 1 | 1.00 |
packaging | n/a | 0.00 | 0.00 | 0 | 0 | 1 | 1 | 1.00 |
risky_diff | n/a | 0.00 | 0.00 | 0 | 0 | 3 | 2 | 1.00 |
secrets | 1.00 | 1.00 | 1.00 | 4 | 0 | 0 | 3 | 1.00 |
sourcemaps | 1.00 | 1.00 | 1.00 | 1 | 0 | 0 | 1 | 1.00 |
| Overall | 1.00 | 0.69 | 0.81 | 11 | 0 | 5 | 10 | 1.00 |
The live regression guard is tests/test_corpus_precision.py, which fails CI if any tp_ case stops firing or any fp_ case starts producing an actionable (HIGH/CRITICAL) finding. Regenerate this report with make bench-precision after adding corpus cases.