confidence-calibration.md

May 14, 2026 Β· View on GitHub

Hack23 Logo

🎯 Confidence Calibration Framework

πŸ“Š Canonical Unified Table: 🟒/🟑/πŸ”΄ Markers Β· WEP Bands Β· Admiralty Grades
🎯 One Table to Rule All Confidence Labels β€” Forward-Looking vs. Evidence-Grounded Claims

Owner Version Effective Date Classification

πŸ“‹ Document Owner: CEO | πŸ“„ Version: 1.0 | πŸ“… Last Updated: 2026-05-14 (UTC) πŸ”„ Review Cycle: Quarterly | ⏰ Next Review: 2026-08-14 🏒 Owner: Hack23 AB (Org.nr 5595347807) | 🏷️ Classification: Public


🎯 Purpose

This document canonicalises all confidence-labelling schemes used across the EU Parliament Monitor analysis library into a single unified reference. It resolves the drift between three schemes that have appeared independently in different runs:

Observed schemeWhere foundCanonical mapping
🟒/🟑/πŸ”΄ tier markersAll artifactsThis document β€” Β§ Tier markers
WEP probability bandsForward-looking claimsThis document β€” Β§ WEP integration
Admiralty A–F Γ— 1–6 gradesSource citationsosint-tradecraft-standards.md Β§2
Tier-1/2/3 data tiersYear-ahead / long-horizon runsThis document β€” Β§ Tier-1/2/3 alignment
Election-cycle WEP bands + Admiralty pairsElectoral artifactsThis document β€” Β§ Electoral domain

After reading this document, all agents must use only the canonical forms defined here. Earlier ad-hoc variants in historical artifacts are noted for backward compatibility only.


πŸ›‘οΈ ISMS Policy Alignment

PolicyRelevance
Hack23 AI_Policy Β§1Objectivity and explained uncertainties β€” confidence labels are the mechanism
Information Security Policy Β§6Evidence handling β€” confidence-in-evidence tracked separately from probability
Secure Development PolicyStructured analytic discipline is the equivalent of SDLC quality gates

1️⃣ The Three Confidence Dimensions

Every analytic claim has up to three independent confidence dimensions. They must be kept separate and never merged into a single score.

%%{init: {"theme":"dark","themeVariables":{"primaryColor":"#1565C0","primaryTextColor":"#ffffff","primaryBorderColor":"#0A3F7F","lineColor":"#90CAF9","secondaryColor":"#2E7D32","secondaryTextColor":"#ffffff","tertiaryColor":"#FF9800","tertiaryTextColor":"#000000","mainBkg":"#1565C0","secondBkg":"#2E7D32","tertiaryBkg":"#FF9800","noteBkgColor":"#FFC107","noteTextColor":"#000000","errorBkgColor":"#D32F2F","fontFamily":"Inter, Helvetica, Arial, sans-serif"}}}%%
graph LR
    A["Analytic claim"] --> D1["Dimension 1<br/>Confidence in EVIDENCE<br/>🟒/🟑/πŸ”΄<br/>(Admiralty grade)"]
    A --> D2["Dimension 2<br/>PROBABILITY of outcome<br/>WEP band<br/>(forward-looking only)"]
    A --> D3["Dimension 3<br/>DATA TIER<br/>Tier 1/2/3<br/>(long-horizon runs)"]

    style D1 fill:#2E7D32,color:#fff
    style D2 fill:#1565C0,color:#fff
    style D3 fill:#7B1FA2,color:#fff
DimensionWhat it expressesApplies toNotation
Confidence in evidenceHow strongly the underlying sources support the claimEvery claim🟒/🟑/πŸ”΄ marker + Admiralty grade
Probability of outcomeHow likely a forward-looking event is to occurForward-looking claims onlyWEP band (Almost Certain β†’ Almost No Chance)
Data tierThe freshness and completeness of the data available this runLong-horizon, electoral, and degraded-mode runsTier 1/2/3 label

Anti-pattern to avoid: Writing "🟑 MEDIUM confidence that the vote will pass with 55–80% probability" conflates evidence confidence (🟑) with probability (WEP band). The correct form is: "Likely (55–80 %) [WEP band] that the vote passes [get_voting_records, A1]. Confidence in evidence: 🟒 HIGH."


2️⃣ Dimension 1 β€” Confidence in Evidence (🟒/🟑/πŸ”΄)

2.1 Canonical Definition

The 🟒/🟑/πŸ”΄ markers express confidence in the underlying evidence β€” whether the sources are reliable and the information credible enough to support the claim. They are derived directly from the Admiralty 6Γ—6 matrix in osint-tradecraft-standards.md Β§2.3.

MarkerLabelAdmiralty combinationsMeaning
🟒HIGHA1, A2, A3, B1, B2, C1Primary-source or tightly corroborated evidence; fit to support headline judgements
🟑MEDIUMA4, A6, B3, B4, B6, C2, C3, D1, D2, F1, F2Indicative but not conclusive; must be flanked by β‰₯1 other piece of evidence before supporting a headline judgement
πŸ”΄LOWA5, B5, C4, C5, C6, D3–D6, E1–E6, F3–F6Noted but never carries a top-level judgement on its own; appears only in caveats, limitations, or monitoring sections

2.2 When to Use Each Marker

🟒 HIGH β€” use when:

  • The claim cites a direct EP MCP record (procedure ID, adopted-text ID, RCV reference).
  • Two independent A/B-grade sources corroborate the same fact.
  • The claim is a formal institutional fact (seat count, treaty majority threshold, committee mandate).

🟑 MEDIUM β€” use when:

  • Only one A-grade source is available without corroboration.
  • The EP MCP feed returned degraded/partial data (see source-triangulation.md Steps 2–3).
  • The claim is structural inference from well-established EP patterns (Step 3 triangulation).
  • Voter/opinion data is based on Eurobarometer proxy rather than current-cycle polling.

πŸ”΄ LOW β€” use when:

  • The claim relies solely on KB integration / institutional-framework analysis (Step 4 triangulation).
  • All primary EP MCP feeds failed and no cross-source triangulation is available.
  • The claim is a tail-risk or wildcard scenario assessment without supporting roll-call data.
  • Electoral projections more than 18 months from the next election.

2.3 Prohibited Inflation

The following upgrades are explicitly prohibited:

Prohibited upgradeCorrect action
Upgrading πŸ”΄ to 🟑 because the analyst "feels confident"Keep πŸ”΄ and disclose the limitation
Omitting the marker because it would be πŸ”΄Include πŸ”΄ with explicit caveat
Using 🟒 for a claim derived from a single C-grade sourceUse 🟑 C2 or πŸ”΄ C3+
Using 🟒 for IMF training-data fallbackUse 🟑 (training-data vintage is at best B2/B3)

3️⃣ Dimension 2 β€” Probability of Outcome (WEP Bands)

3.1 Canonical WEP Bands

The WEP bands are defined in osint-tradecraft-standards.md Β§3. The canonical form reproduced here for reference:

BandPhraseNumeric rangeWhen to use
1Almost no chance / remote1–5 %Tail risk; PfE–EPP joint report on Rule-of-Law
2Very unlikely / highly improbable5–20 %Grand-Coalition rupture before 2026 mid-term
3Unlikely / improbable20–45 %Unilateral Council blocking
4Roughly even chance45–55 %ECR whip decision on migration
5Likely / probable55–80 %EPP–Renew compromise survives first reading
6Very likely / highly probable80–95 %Budget discharge on first vote
7Almost certain / nearly certain95–99 %Interinstitutional agreement signature

3.2 WEP and Evidence-Confidence Are Independent

The table below illustrates that a claim can have any combination of WEP band and evidence confidence:

WEP bandEvidence confidenceCanonical formMeaning
Likely (55–80 %)🟒 HIGH[A1] Likely (55–80 %)Strong evidence; claim is probably true
Likely (55–80 %)🟑 MEDIUM[B3] Likely (55–80 %)Weak evidence but structural inference supports
Almost certain (95–99 %)πŸ”΄ LOW[F6] Almost certain (95–99 %)Institutional convention; no primary data this run
Very unlikely (5–20 %)🟒 HIGH[A1] Very unlikely (5–20 %)Strong evidence that outcome is unlikely

3.3 Canonical Notation Form

Every forward-looking claim must follow this exact form:

"[WEP phrase] ([numeric %]) [Admiralty grade]. Confidence in evidence: [🟒/🟑/πŸ”΄] [level]."

Example:

"A Grand-Coalition majority on the migration-pact amendments is likely (55–80 %) [get_voting_records + historical-baseline, A1]. Confidence in evidence: 🟒 HIGH."

3.4 Time Horizon Requirement

Every WEP-banded claim names an explicit horizon (see osint-tradecraft-standards.md Β§3.4):

HorizonWindowExample
Tactical0–7 days"…before the May plenary (WEP: likely, A1, tactical)"
Operational7–90 days"…within this session (WEP: likely, A2, operational)"
Strategic90 days – 18 months"…within this term (WEP: possible, B3, strategic)"
Structural18+ months"…across multiple terms (WEP: unlikely, C3, structural)"

4️⃣ Dimension 3 β€” Data Tier (Long-Horizon Runs)

Long-horizon article types (year-ahead, term-outlook, election-cycle) use a three-tier data classification to convey how fresh and primary the underlying data is. The tier is an input to confidence selection, not a replacement for it.

4.1 Tier Definitions

TierLabelMeaningMaps to confidence
Tier 1Primary/current dataEP MCP data fetched this run (< 7 days old); IMF SDMX live response; Eurostat live queryEligible for 🟒 HIGH (subject to Admiralty grade)
Tier 2Recent secondary dataData 7–90 days old; feed data with FRESHNESS_FALLBACK; IMF training-data vintage (recent)Maximum 🟑 MEDIUM
Tier 3Structural / historical dataData > 90 days old; prior-term datasets; academic historical series; KB-only institutional factsMaximum πŸ”΄ LOW unless the claim is a formal institutional fact (e.g. seat counts, RoP majorities)

4.2 Tier Label in Artifacts

Long-horizon artifacts add a Data tier: [Tier N] tag in the artifact header and in each major section where the tier changes. Example:

## 3. Forecast: 2027–2029 Coalition Landscape

> **Data tier:** Tier 2 (Eurobarometer Q4 2025 + EP10 historical RCV patterns)
> **Max confidence:** 🟑 MEDIUM

4.3 Tier Reduction Factors

The reference-quality-thresholds.json schema already applies line-floor reduction factors for degraded modes. The tier system maps to these factors:

Mode in manifest.jsonTier alignmentFloor reduction
fullAll Tier 11.0Γ—
title-onlyTier 2 (partial)0.75Γ—
degraded-imfTier 2 (economic)0.85Γ—
degraded-votingTier 2 (coalition)0.85Γ—
minimalTier 3 dominant0.65Γ—

5️⃣ Electoral Domain Alignment

Electoral artifacts (election-cycle, seat-projection, voter-segmentation) use a combined WEP + Admiralty form that pairs electoral probability with source grade. The canonical form:

"EPP seat share at EP2029: [WEP band] to hold [N–M seats] ([%] of total). Source: [Eurobarometer Q4 2025, B2] + [EP2024 election turnout by MS, A1]. Confidence: 🟑 MEDIUM."

The following anti-patterns are specifically forbidden in electoral artifacts:

Anti-patternCorrection
"Predicted 192 EPP seats (HIGH confidence)" without WEP + AdmiraltyAdd WEP band + source grade
"WEP: likely β€” EPP gains seats" without evidence basisState evidence basis and Admiralty grade
Mixing πŸ”΄ LOW confidence Tier-3 structural forecasts with 🟒 HIGH headlineSeparate claims by tier; use 🟑 MEDIUM ceiling

6️⃣ Worked Examples

6.1 β€” Evidence-based political judgement (standard form)

Claim: EPP group cohesion on the Banking Union vote was strong. Evidence: get_voting_records returned RCV-2026-0412 showing 181/188 EPP MEPs voted in line with group position. Admiralty: A1 (direct plenary record, single EP source, uncontested). Confidence in evidence: 🟒 HIGH. No WEP needed β€” this is a retrospective fact, not a forward probability.

Canonical form: "EPP group cohesion on Banking Union was 96.3 % (181/188 MEPs) [get_voting_records, RCV-2026-0412, A1]. Confidence: 🟒 HIGH."

6.2 β€” Forward-looking claim on coalition behaviour

Claim: The Grand Coalition is likely to hold together on the upcoming AI-Act vote. Evidence: analyze_coalition_dynamics + historical-baseline (EPP–S&D–Renew agreement rate 84 % over trailing 30 days, A2). WEP: Likely (55–80 %) β€” tactical horizon (next 14 days). Confidence in evidence: 🟒 HIGH.

Canonical form: "The Grand Coalition on AI-Act is likely to hold (55–80 %, tactical) [analyze_coalition_dynamics + historical-baseline, A2]. Confidence in evidence: 🟒 HIGH."

6.3 β€” Degraded-data claim (Step 3 triangulation)

Claim: PfE group cohesion is structurally low due to national-party heterogeneity. Evidence: EP MCP roll-call feed unavailable; inference from Hooghe/Marks EU-attitudes literature (C2) + ECR/PfE historical voting patterns from prior-term data (B3). Triangulation step: 3 (structural inference). Confidence in evidence: 🟑 MEDIUM (Step 3 ceiling).

Canonical form: "PfE group cohesion is structurally low [TRIANGULATION-STEP: 3 β€” structural inference from Hooghe/Marks (C2) + prior-term patterns (B3)]. Confidence: 🟑 MEDIUM."

6.4 β€” Long-horizon electoral forecast

Claim: EPP is likely to retain the largest group share at EP2029. Evidence: Eurobarometer Q4 2025 (B2) + EP10 seat distribution (A1) + structural incumbency advantage literature (C3). Tier: Tier 2 (secondary data). WEP: Likely (55–80 %) β€” structural horizon. Confidence in evidence: 🟑 MEDIUM (Tier 2 ceiling).

Canonical form: "EPP is likely (55–80 %, structural) to retain the largest group at EP2029 [Eurobarometer Q4 2025 B2 + EP10 seat distribution A1]. Data tier: 2. Confidence: 🟑 MEDIUM."


πŸ“‹ Quick-Reference Checklist

  • Three dimensions separate: Evidence confidence, probability (WEP), and data tier are never merged.
  • Canonical notation: Every forward-looking claim uses [WEP phrase] ([%]) [Admiralty grade]. Confidence: [🟒/🟑/πŸ”΄].
  • Time horizon present: Every WEP claim includes tactical/operational/strategic/structural horizon.
  • No inflation: 🟒 never used for Step-3/4 triangulation claims; IMF training-data is 🟑 maximum.
  • Electoral domain: WEP + Admiralty pair used in all electoral artifact probability claims.
  • Tier label in long-horizon artifacts: Data tier: [N] in every major section of year-ahead, term-outlook, election-cycle artifacts.


Document Control:

  • Path: analysis/methodologies/confidence-calibration.md
  • Classification: Public
  • Version: 1.0 β€” Initial unified table resolving 🟒/🟑/πŸ”΄ + WEP + Admiralty + Tier-1/2/3 drift observed across 2026-05-14 run set reflections.
  • ISMS Reference: Hack23/ISMS-PUBLIC AI Policy Β§1 (Objectivity and explained uncertainties)