πŸ“ Reference Analysis Quality Template

May 3, 2026 Β· View on GitHub

πŸ“Œ Template Instructions: Copy to analysis/daily/{date}/{article-type}-run{N}/intelligence/reference-analysis-quality.md. Self-score this run against the reference benchmark (Run 184, 2026-04-18) plus Pass-2 action list. See methodologies/per-artifact-methodologies.md Β§reference-analysis-quality.

🎯 Purpose: Mandatory quality gate comparing current run to the gold-standard reference. Identifies gaps and provides concrete Pass-2 improvement targets.


πŸ“‹ Document Metadata

FieldValue
Report ID[REQUIRED: RAQ-YYYY-MM-DD-runNN]
Current Run[REQUIRED: {type}-run{N}, YYYY-MM-DD]
Reference Benchmark[REQUIRED: breaking-run184, 2026-04-18]
Pass Number[REQUIRED: Pass 1 / Pass 2]
Benchmark Met?[REQUIRED: βœ… Yes / ⚠️ Partial / ❌ No]

1️⃣ Per-Artifact Line Count vs. Depth Floor

Artifact (run-relative path)Depth FloorActual LinesDeltaStatus
intelligence/synthesis-summary.md205[#][Β±#][βœ…/⚠️/❌]
intelligence/analysis-index.md160[#][Β±#][βœ…/⚠️/❌]
intelligence/voting-patterns.md150[#][Β±#][βœ…/⚠️/❌]
intelligence/coalition-dynamics.md135[#][Β±#][βœ…/⚠️/❌]
intelligence/stakeholder-map.md305[#][Β±#][βœ…/⚠️/❌]
intelligence/scenario-forecast.md280[#][Β±#][βœ…/⚠️/❌]
intelligence/pestle-analysis.md250[#][Β±#][βœ…/⚠️/❌]
intelligence/threat-model.md250[#][Β±#][βœ…/⚠️/❌]
intelligence/economic-context.md185[#][Β±#][βœ…/⚠️/❌]
intelligence/historical-baseline.md190[#][Β±#][βœ…/⚠️/❌]
intelligence/mcp-reliability-audit.md385[#][Β±#][βœ…/⚠️/❌]
risk-scoring/risk-matrix.md150[#][Β±#][βœ…/⚠️/❌]
risk-scoring/quantitative-swot.md140[#][Β±#][βœ…/⚠️/❌]
intelligence/cross-run-diff.md100[#][Β±#][βœ…/⚠️/❌]
intelligence/workflow-audit.md100[#][Β±#][βœ…/⚠️/❌]

Note: Floors above are the breaking defaults. For other article types, consult analysis/methodologies/reference-quality-thresholds.json β€” the validator keys floors by run-relative path (e.g. intelligence/synthesis-summary.md, risk-scoring/risk-matrix.md).

Status definitions:

  • βœ… At or above floor
  • ⚠️ Within 10% of floor (acceptable with LOW confidence marker)
  • ❌ Below 90% of floor (breaking failure)

Artifacts below floor: [REQUIRED: count or "none"]


2️⃣ Per-Artifact Mermaid Presence & Theme Compliance

ArtifactMermaid Required?Count PresentTheme Init Block?Status
synthesis-summary.mdYes (β‰₯2)[#][βœ…/❌][βœ…/⚠️/❌]
voting-patterns.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
coalition-dynamics.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
stakeholder-map.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
scenario-forecast.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
pestle-analysis.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
threat-model.mdYes (β‰₯2)[#][βœ…/❌][βœ…/⚠️/❌]
risk-matrix.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
quantitative-swot.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
historical-baseline.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]
wildcards-blackswans.mdYes (β‰₯1)[#][βœ…/❌][βœ…/⚠️/❌]

Mermaid failures: [REQUIRED: list artifacts missing required diagrams or missing theme init block]

Theme compliance check:

%%{init: {"theme":"dark","themeVariables":{...}}}%%

[REQUIRED: βœ… All diagrams use standard theme block / ❌ Some diagrams missing theme]


3️⃣ Per-Artifact Evidence-Density Score

Methodology: Count citations per 100 lines (procedure IDs, RCV IDs, document IDs, indicator codes, MEP names with roles).

ArtifactLinesCitationsCitations/100LBenchmark TargetStatus
synthesis-summary.md[#][#][#.#]β‰₯3.0[βœ…/⚠️/❌]
voting-patterns.md[#][#][#.#]β‰₯3.5[βœ…/⚠️/❌]
stakeholder-map.md[#][#][#.#]β‰₯4.0[βœ…/⚠️/❌]
scenario-forecast.md[#][#][#.#]β‰₯2.5[βœ…/⚠️/❌]
threat-model.md[#][#][#.#]β‰₯3.0[βœ…/⚠️/❌]
economic-context.md[#][#][#.#]β‰₯5.0[βœ…/⚠️/❌]
risk-matrix.md[#][#][#.#]β‰₯2.0[βœ…/⚠️/❌]

Artifacts below benchmark: [REQUIRED: count or "none"]


4️⃣ Benchmark Gap Narrative

`[REQUIRED: β‰₯100 words identifying the run's weakest artifacts and why they fall short of reference quality. Examples:

  • "stakeholder-map.md at 260 lines (floor 305) β€” quadrant narratives are 80 words each instead of required β‰₯150"
  • "economic-context.md missing IMF forward projections β€” cites only backward-looking historical data; IMF is the mandatory primary source"
  • "threat-model.md attack tree has only 2 levels instead of required β‰₯3"
  • "synthesis-summary.md forward monitors lack date-bounded trigger events"

For each gap, cite the specific methodology requirement and the reference-run example that demonstrates compliance.]`

Gap summary table:

ArtifactGap TypeSeverityReference Example
[REQUIRED: artifact name][REQUIRED: Line count / Evidence / Mermaid / Section depth][🟒 Minor / 🟑 Moderate / πŸ”΄ Breaking][REQUIRED: cite reference-run section or line]
[REQUIRED][REQUIRED][...][REQUIRED]
[REQUIRED][REQUIRED][...][REQUIRED]

5️⃣ Pass-2 Action List

%%{init: {"theme":"dark","themeVariables":{"primaryColor":"#1565C0","primaryTextColor":"#ffffff","primaryBorderColor":"#0A3F7F","lineColor":"#90CAF9","secondaryColor":"#2E7D32","secondaryTextColor":"#ffffff","tertiaryColor":"#FF9800","tertiaryTextColor":"#000000","mainBkg":"#1565C0","secondBkg":"#2E7D32","tertiaryBkg":"#FF9800","noteBkgColor":"#FFC107","noteTextColor":"#000000","errorBkgColor":"#D32F2F","fontFamily":"Inter, Helvetica, Arial, sans-serif"}}}%%
flowchart LR
    PASS1[Pass 1<br/>Complete] -->|gaps detected| GAP{Benchmark<br/>Gap Analysis}
    GAP -->|action list| PASS2[Pass 2<br/>Improvement]
    PASS2 -->|re-validate| REF[Reference-Quality<br/>Exit]
    
    style PASS1 fill:#1565C0,color:#ffffff
    style GAP fill:#FF9800,color:#000000
    style PASS2 fill:#7B1FA2,color:#ffffff
    style REF fill:#2E7D32,color:#ffffff

Specific actions for Pass 2:

  1. [REQUIRED: Artifact name] Β· [REQUIRED: Section name]
    Action: [REQUIRED: β‰₯40 words describing exactly what to add, expand, or revise. Must be concrete enough that Pass 2 can execute without re-reading the entire file. Examples: "Expand stakeholder-map.md Β§4.2 Champions quadrant from 80 to β‰₯150 words by adding named MEPs with committee roles", "Add missing IMF WEO projection table to economic-context.md Β§4 covering 2025-2029".]

  2. [REQUIRED: Artifact name] Β· [REQUIRED: Section name]
    Action: [REQUIRED: β‰₯40 words]

  3. [REQUIRED]
    Action: [REQUIRED]

  4. [REQUIRED]
    Action: [REQUIRED]

  5. [REQUIRED]
    Action: [REQUIRED]

Total actions: [REQUIRED: count β€” typically 5-15 for a Pass-1 run]

Estimated Pass-2 time: [REQUIRED: HH:MM based on action complexity]


6️⃣ Reference-Run Comparison

Run 184 metrics (benchmark):

MetricRun 184This RunDelta
Total artifacts produced18[#][Β±#]
Artifacts meeting depth floor18 (100%)[#][Β±#]
Total line count (all artifacts)4,872[#][Β±#]
Mermaid diagrams26[#][Β±#]
Evidence citations187[#][Β±#]
Average citations/100L3.84[#.#][Β±#.#]

What Run 184 did exceptionally well:

[REQUIRED: β‰₯80 words highlighting 2-3 specific strengths of the reference run that this run should emulate. Examples: "Run 184 stakeholder-map.md included movement-since-prior-period analysis with β‰₯5 actors tracked across runs", "Run 184 synthesis-summary.md had 8 forward monitors with specific date triggers, not generic 'watch for...' statements".]


7️⃣ Validator Output Integration

Automated validator findings:

[REQUIRED: If an automated validator was run (checking depth floors, Mermaid presence, citation counts), paste or summarize its output here. If no validator was run, note "manual validation only".]

Validator failures addressed: [REQUIRED: count or "none"]
Validator warnings acknowledged: [REQUIRED: count or "none"]


8️⃣ Confidence Assessment

Overall reference-quality confidence: [REQUIRED: 🟒 HIGH / 🟑 MEDIUM / πŸ”΄ LOW]

Confidence rationale: [REQUIRED: β‰₯60 words explaining whether this run meets, approaches, or falls short of reference quality. If LOW, what are the 2-3 most severe gaps?]


πŸ› οΈ Worked example β€” reference-quality scoring for a hypothetical run

Quality dimensionRun scoreReference benchmarkGapAction
Word count vs floor6 200 / 5 000124%+24%Pass
Procedures cited inline18β‰₯15+3Pass
Mermaid diagrams7β‰₯3+4Pass
Historical comparisons2β‰₯20Pass
Coalition-cohesion citations9β‰₯5+4Pass
IMF + WB combined sources4 IMF + 2 WBβ‰₯3 IMF+3 IMF, +2 WBPass
Admiralty grades attached91%β‰₯80%+11 ppPass
WEP bands on forecasts100%100%0Pass
Pass-2 expansion ratio1.65Γ—β‰₯1.5Γ—+0.15Γ—Pass
Pre/post HTML clean diff000Pass

Aggregate: 10/10 dimensions pass; reference quality met. Gap area to monitor: historical-comparison count is at floor β€” adding 1 more multi-year baseline would strengthen rigour.

🚫 Anti-patterns β€” reference-quality-quality failures

Anti-patternWhy it failsCorrect approach
Self-attestation onlySubjectiveEach dimension has measurable threshold
Score without comparisonNo anchorCompare to a named benchmark run
Pass-2 ratio < 1.5Γ—Skipped iterationExpand thin sections in genuine Pass-2
"Reference quality" without dimensionsUnverifiableList all measurable dimensions
Run validator failures hiddenStage-C bypassAcknowledge any validator output
Pass without rationaleAudit trail gapEach dimension cite the count

🎯 EP MCP tool inputs

This artifact is introspective β€” it reports on the artifact set and article rather than calling additional MCP tools. Inputs:

  • The run's full artifact set + article
  • Prior reference-benchmark runs (e.g. Run 184)
  • Validator output (if any)

πŸ”— Controlling methodology cross-references

βœ… Stage-C completeness signals

  • Line floor: 190 lines
  • β‰₯ 10 quality dimensions scored
  • Each dimension: actual / benchmark / gap / verdict
  • Reference benchmark run cited
  • Validator output acknowledged (or "no validator failures")

Document Control: /analysis/daily/{date}/{type}-run{N}/intelligence/reference-analysis-quality.md Β· Template v1.2 Β· Depth floor: 190 lines.