π Reference Analysis Quality Template
May 3, 2026 Β· View on GitHub
π Template Instructions: Copy to
analysis/daily/{date}/{article-type}-run{N}/intelligence/reference-analysis-quality.md. Self-score this run against the reference benchmark (Run 184, 2026-04-18) plus Pass-2 action list. See methodologies/per-artifact-methodologies.md Β§reference-analysis-quality.
π― Purpose: Mandatory quality gate comparing current run to the gold-standard reference. Identifies gaps and provides concrete Pass-2 improvement targets.
π Document Metadata
| Field | Value |
|---|---|
| Report ID | [REQUIRED: RAQ-YYYY-MM-DD-runNN] |
| Current Run | [REQUIRED: {type}-run{N}, YYYY-MM-DD] |
| Reference Benchmark | [REQUIRED: breaking-run184, 2026-04-18] |
| Pass Number | [REQUIRED: Pass 1 / Pass 2] |
| Benchmark Met? | [REQUIRED: β
Yes / β οΈ Partial / β No] |
1οΈβ£ Per-Artifact Line Count vs. Depth Floor
| Artifact (run-relative path) | Depth Floor | Actual Lines | Delta | Status |
|---|---|---|---|---|
intelligence/synthesis-summary.md | 205 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/analysis-index.md | 160 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/voting-patterns.md | 150 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/coalition-dynamics.md | 135 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/stakeholder-map.md | 305 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/scenario-forecast.md | 280 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/pestle-analysis.md | 250 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/threat-model.md | 250 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/economic-context.md | 185 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/historical-baseline.md | 190 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/mcp-reliability-audit.md | 385 | [#] | [Β±#] | [β
/β οΈ/β] |
risk-scoring/risk-matrix.md | 150 | [#] | [Β±#] | [β
/β οΈ/β] |
risk-scoring/quantitative-swot.md | 140 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/cross-run-diff.md | 100 | [#] | [Β±#] | [β
/β οΈ/β] |
intelligence/workflow-audit.md | 100 | [#] | [Β±#] | [β
/β οΈ/β] |
Note: Floors above are the
breakingdefaults. For other article types, consultanalysis/methodologies/reference-quality-thresholds.jsonβ the validator keys floors by run-relative path (e.g.intelligence/synthesis-summary.md,risk-scoring/risk-matrix.md).
Status definitions:
- β At or above floor
- β οΈ Within 10% of floor (acceptable with LOW confidence marker)
- β Below 90% of floor (breaking failure)
Artifacts below floor: [REQUIRED: count or "none"]
2οΈβ£ Per-Artifact Mermaid Presence & Theme Compliance
| Artifact | Mermaid Required? | Count Present | Theme Init Block? | Status |
|---|---|---|---|---|
synthesis-summary.md | Yes (β₯2) | [#] | [β
/β] | [β
/β οΈ/β] |
voting-patterns.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
coalition-dynamics.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
stakeholder-map.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
scenario-forecast.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
pestle-analysis.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
threat-model.md | Yes (β₯2) | [#] | [β
/β] | [β
/β οΈ/β] |
risk-matrix.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
quantitative-swot.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
historical-baseline.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
wildcards-blackswans.md | Yes (β₯1) | [#] | [β
/β] | [β
/β οΈ/β] |
Mermaid failures: [REQUIRED: list artifacts missing required diagrams or missing theme init block]
Theme compliance check:
%%{init: {"theme":"dark","themeVariables":{...}}}%%
[REQUIRED: β
All diagrams use standard theme block / β Some diagrams missing theme]
3οΈβ£ Per-Artifact Evidence-Density Score
Methodology: Count citations per 100 lines (procedure IDs, RCV IDs, document IDs, indicator codes, MEP names with roles).
| Artifact | Lines | Citations | Citations/100L | Benchmark Target | Status |
|---|---|---|---|---|---|
synthesis-summary.md | [#] | [#] | [#.#] | β₯3.0 | [β
/β οΈ/β] |
voting-patterns.md | [#] | [#] | [#.#] | β₯3.5 | [β
/β οΈ/β] |
stakeholder-map.md | [#] | [#] | [#.#] | β₯4.0 | [β
/β οΈ/β] |
scenario-forecast.md | [#] | [#] | [#.#] | β₯2.5 | [β
/β οΈ/β] |
threat-model.md | [#] | [#] | [#.#] | β₯3.0 | [β
/β οΈ/β] |
economic-context.md | [#] | [#] | [#.#] | β₯5.0 | [β
/β οΈ/β] |
risk-matrix.md | [#] | [#] | [#.#] | β₯2.0 | [β
/β οΈ/β] |
Artifacts below benchmark: [REQUIRED: count or "none"]
4οΈβ£ Benchmark Gap Narrative
`[REQUIRED: β₯100 words identifying the run's weakest artifacts and why they fall short of reference quality. Examples:
- "stakeholder-map.md at 260 lines (floor 305) β quadrant narratives are 80 words each instead of required β₯150"
- "economic-context.md missing IMF forward projections β cites only backward-looking historical data; IMF is the mandatory primary source"
- "threat-model.md attack tree has only 2 levels instead of required β₯3"
- "synthesis-summary.md forward monitors lack date-bounded trigger events"
For each gap, cite the specific methodology requirement and the reference-run example that demonstrates compliance.]`
Gap summary table:
| Artifact | Gap Type | Severity | Reference Example |
|---|---|---|---|
[REQUIRED: artifact name] | [REQUIRED: Line count / Evidence / Mermaid / Section depth] | [π’ Minor / π‘ Moderate / π΄ Breaking] | [REQUIRED: cite reference-run section or line] |
[REQUIRED] | [REQUIRED] | [...] | [REQUIRED] |
[REQUIRED] | [REQUIRED] | [...] | [REQUIRED] |
5οΈβ£ Pass-2 Action List
%%{init: {"theme":"dark","themeVariables":{"primaryColor":"#1565C0","primaryTextColor":"#ffffff","primaryBorderColor":"#0A3F7F","lineColor":"#90CAF9","secondaryColor":"#2E7D32","secondaryTextColor":"#ffffff","tertiaryColor":"#FF9800","tertiaryTextColor":"#000000","mainBkg":"#1565C0","secondBkg":"#2E7D32","tertiaryBkg":"#FF9800","noteBkgColor":"#FFC107","noteTextColor":"#000000","errorBkgColor":"#D32F2F","fontFamily":"Inter, Helvetica, Arial, sans-serif"}}}%%
flowchart LR
PASS1[Pass 1<br/>Complete] -->|gaps detected| GAP{Benchmark<br/>Gap Analysis}
GAP -->|action list| PASS2[Pass 2<br/>Improvement]
PASS2 -->|re-validate| REF[Reference-Quality<br/>Exit]
style PASS1 fill:#1565C0,color:#ffffff
style GAP fill:#FF9800,color:#000000
style PASS2 fill:#7B1FA2,color:#ffffff
style REF fill:#2E7D32,color:#ffffff
Specific actions for Pass 2:
-
[REQUIRED: Artifact name]Β·[REQUIRED: Section name]
Action:[REQUIRED: β₯40 words describing exactly what to add, expand, or revise. Must be concrete enough that Pass 2 can execute without re-reading the entire file. Examples: "Expand stakeholder-map.md Β§4.2 Champions quadrant from 80 to β₯150 words by adding named MEPs with committee roles", "Add missing IMF WEO projection table to economic-context.md Β§4 covering 2025-2029".] -
[REQUIRED: Artifact name]Β·[REQUIRED: Section name]
Action:[REQUIRED: β₯40 words] -
[REQUIRED]
Action:[REQUIRED] -
[REQUIRED]
Action:[REQUIRED] -
[REQUIRED]
Action:[REQUIRED]
Total actions: [REQUIRED: count β typically 5-15 for a Pass-1 run]
Estimated Pass-2 time: [REQUIRED: HH:MM based on action complexity]
6οΈβ£ Reference-Run Comparison
Run 184 metrics (benchmark):
| Metric | Run 184 | This Run | Delta |
|---|---|---|---|
| Total artifacts produced | 18 | [#] | [Β±#] |
| Artifacts meeting depth floor | 18 (100%) | [#] | [Β±#] |
| Total line count (all artifacts) | 4,872 | [#] | [Β±#] |
| Mermaid diagrams | 26 | [#] | [Β±#] |
| Evidence citations | 187 | [#] | [Β±#] |
| Average citations/100L | 3.84 | [#.#] | [Β±#.#] |
What Run 184 did exceptionally well:
[REQUIRED: β₯80 words highlighting 2-3 specific strengths of the reference run that this run should emulate. Examples: "Run 184 stakeholder-map.md included movement-since-prior-period analysis with β₯5 actors tracked across runs", "Run 184 synthesis-summary.md had 8 forward monitors with specific date triggers, not generic 'watch for...' statements".]
7οΈβ£ Validator Output Integration
Automated validator findings:
[REQUIRED: If an automated validator was run (checking depth floors, Mermaid presence, citation counts), paste or summarize its output here. If no validator was run, note "manual validation only".]
Validator failures addressed: [REQUIRED: count or "none"]
Validator warnings acknowledged: [REQUIRED: count or "none"]
8οΈβ£ Confidence Assessment
Overall reference-quality confidence: [REQUIRED: π’ HIGH / π‘ MEDIUM / π΄ LOW]
Confidence rationale: [REQUIRED: β₯60 words explaining whether this run meets, approaches, or falls short of reference quality. If LOW, what are the 2-3 most severe gaps?]
π οΈ Worked example β reference-quality scoring for a hypothetical run
| Quality dimension | Run score | Reference benchmark | Gap | Action |
|---|---|---|---|---|
| Word count vs floor | 6 200 / 5 000 | 124% | +24% | Pass |
| Procedures cited inline | 18 | β₯15 | +3 | Pass |
| Mermaid diagrams | 7 | β₯3 | +4 | Pass |
| Historical comparisons | 2 | β₯2 | 0 | Pass |
| Coalition-cohesion citations | 9 | β₯5 | +4 | Pass |
| IMF + WB combined sources | 4 IMF + 2 WB | β₯3 IMF | +3 IMF, +2 WB | Pass |
| Admiralty grades attached | 91% | β₯80% | +11 pp | Pass |
| WEP bands on forecasts | 100% | 100% | 0 | Pass |
| Pass-2 expansion ratio | 1.65Γ | β₯1.5Γ | +0.15Γ | Pass |
| Pre/post HTML clean diff | 0 | 0 | 0 | Pass |
Aggregate: 10/10 dimensions pass; reference quality met. Gap area to monitor: historical-comparison count is at floor β adding 1 more multi-year baseline would strengthen rigour.
π« Anti-patterns β reference-quality-quality failures
| Anti-pattern | Why it fails | Correct approach |
|---|---|---|
| Self-attestation only | Subjective | Each dimension has measurable threshold |
| Score without comparison | No anchor | Compare to a named benchmark run |
| Pass-2 ratio < 1.5Γ | Skipped iteration | Expand thin sections in genuine Pass-2 |
| "Reference quality" without dimensions | Unverifiable | List all measurable dimensions |
| Run validator failures hidden | Stage-C bypass | Acknowledge any validator output |
| Pass without rationale | Audit trail gap | Each dimension cite the count |
π― EP MCP tool inputs
This artifact is introspective β it reports on the artifact set and article rather than calling additional MCP tools. Inputs:
- The run's full artifact set + article
- Prior reference-benchmark runs (e.g. Run 184)
- Validator output (if any)
π Controlling methodology cross-references
../methodologies/ai-driven-analysis-guide.md Β§Step 10 / 10.5../methodologies/per-artifact-methodologies.md Β§reference-analysis-qualitymethodology-reflection.mdβ the 10.5 reflection artifact
β Stage-C completeness signals
- Line floor: 190 lines
- β₯ 10 quality dimensions scored
- Each dimension: actual / benchmark / gap / verdict
- Reference benchmark run cited
- Validator output acknowledged (or "no validator failures")
Document Control: /analysis/daily/{date}/{type}-run{N}/intelligence/reference-analysis-quality.md Β· Template v1.2 Β· Depth floor: 190 lines.