06
July 10, 2026 · View on GitHub
This document extends the constitution (docs/01-foundations.md P6–P7) and the affective layers (docs/05) with the layer the founder asked for in v0.5: when should Engram reach for an interactive visual, for whom, and under what discipline? It exists because of a founder memory — years of "features/dimensions" refusing to click from prose, then one draggable face (each slider a feature, the face morphing live) and the concept landing in seconds — and because of the founder's own caution about it: "this might be just me."
That caution is the right instinct, and this document is the audit. The claims below were assembled the same way docs/05 was built: a fan-out research pass (5 search angles, 27 primary sources fetched, 135 claims extracted; the 25 load-bearing ones each adversarially verified by three independent refute-first voters — 23 survived 3-0, 2 were killed and are listed here as do-not-build-on). Effect sizes are meta-analytic wherever one exists; where a number comes from a single lab or a single study, that is said in place. Four question areas produced nothing that survived verification; they are listed as open questions, and every v0.5 design choice touching them is deliberately conservative.
The verdict, stated up front so the rest can be checked against it:
The interactive visual is a real but conditional medium — and the conditions, not the medium, carry most of the effect. Manipulable models are the strongest interactivity result in the verified base (simulations g+ = 0.62), but guidance inside the artifact is the active ingredient (scaffolded versions beat identical unscaffolded ones, g+ = 0.60; learner control per se is worth g = 0.05 ≈ nothing), the payoff concentrates where the dynamics are the content (representational d = 0.40 vs decorational ≈ −0.05; procedural-motor d = 1.06), and every effect is bounded by expertise reversal. So Engram's explorables stay contract-bound, become content-triggered (the node declares its visual affordance, per Willingham's rule) and learner-dialed (a preference setting, honored as motivation, measured as evidence) — and never become decoration, sandboxes, or a "visual learner" accommodation.
Pillar 15 — The guided manipulable: interactivity is spent on cognition, never on navigation
Claim. A manipulable model of a concept's causal structure, wrapped in prediction gates, scaffolds, and self-explanation prompts, is an evidence-backed encoding medium — because of the wrapping. Strip the guidance and the medium's advantage collapses; add decoration and it reverses.
Evidence — when the medium helps.
- Interactive simulations (the learner manipulates parameters of an underlying model — exactly Engram's Contract clause 2) beat non-simulation instruction at g+ = 0.62 (D'Angelo et al./SRI 2014, k = 46, CI 0.45–0.79; robust to design type and publication-bias checks; honesty flags: a Gates-funded technical report later journal-condensed, and heterogeneity is high, I² ≈ 79). Rutten, van Joolingen & van der Veen (2012) independently conclude "robust evidence that computer simulations can enhance traditional instruction." This is the largest interactivity effect that survived verification.
- Dynamic beats static, modestly, when the motion is the message. The field estimate shrank as the corpus grew: Höffler & Leutner (2007) d = 0.37 → Berney & Bétrancourt (2016; 61 experiments, N = 7,036) g = 0.226. The moderators do the real work: representational animation (the dynamics are the to-be-learned content) d = 0.40 vs decorational ≈ −0.05; procedural-motor content d = 1.06 (the human-movement effect, the largest moderator in the corpus).
- The multimedia corpus is real but conditional. Across Mayer's entire coded corpus 1990–2022 (Cromley & Chen 2025: 92 articles, 591 effects), the principles average g = 0.37, significantly moderated by everything they coded — including a decline per publication year. Treat every principle as conditional, never as law. The within-corpus priority order (single-lab numbers, larger than field-wide): remove seductive details g = 1.00, modality 0.82, personalization 0.70, multimedia proper 0.68, coherence 0.63, self-explanation 0.46, testing 0.41, scaffolding 0.38, embodiment 0.35, cueing 0.24.
Evidence — where it collapses (the leash, at least as robust as the license).
- Guidance is the active ingredient. Adding instructional enhancements to an otherwise-identical simulation: g+ = 0.49; scaffolding alone g+ = 0.60 (SRI 2014; independently: guidance in inquiry d = 0.50, Lazonder & Harmsen 2016). Unassisted discovery underperforms explicit instruction (Alfieri et al. 2011: d = −0.38), while enhanced/guided discovery beats other methods (d = +0.30) — the explorable-with-prediction-gates format is precisely the winning cell (Kirschner, Sweller & Clark 2006; their "minimal guidance" framing is contested for PBL/inquiry, but the unguided-discovery core is conceded even by their critics; productive failure, deliberately sequenced exploration-then-instruction, is the recognized carve-out and is already Engram's PREDICT→STRUGGLE→RESOLVE grammar).
- Learner control per se is worth nothing: g = 0.05 across educational technology (Karich, Burns & Maki 2014; consistent with Niemiec 1996 and Landers & Reddock 2017). The one qualified carve-out: segmentation — learner-advanced segment boundaries help (d ≈ 0.42, Rey et al. 2019) but system-triggered pausing works at least as well as learner-triggered. And the animation advantage itself held only under system pacing (B&B 2016: g = 0.309; with learner playback control, not significant). So: the learner advances between segments; within a segment the dynamic runs itself; scrub bars and navigation freedom are not learning features.
- Concurrent on-screen text kills the animation advantage (B&B 2016: no-text g = 0.883, narration g = 0.336, written-text-during-motion n.s.) — with the verified scope warning that the narration-over-text rule is a system-paced phenomenon and can reverse under learner-paced conditions. Engram's rule: text before the motion (the prediction) and after it (the explanation), never over it.
- Seductive details reliably hurt: g ≈ −0.33 (Sundararajan & Adesope 2020, 58 studies) to −0.16 (2025 multilevel re-analysis) — real in direction, smaller than folklore, and the reason the Contract's "zero decoration" clause outranks every clever idea.
- Expertise reversal is confirmed and disordinal (Tetzlaff et al. 2025: 176 effects, N = 5,924): low-prior-knowledge learners gain from high assistance d = +0.505, high-prior-knowledge learners are harmed by it d = −0.428 (study-level difference d = 0.971; heterogeneity high, I² ≈ 88–91%; "experts" here means more-knowledgeable novices). The meta-analysts' own instruction, adopted verbatim as Engram's default: "rather provide assistance than to withhold it when in doubt." And the sting for this document specifically: interactivity itself shows the reversal — manipulable pictures were optimal for high-prerequisite learners and imposed extraneous load on novices (Schnotz & Rasch 2005; single study, n = 13/group, unreplicated — flagged medium-confidence, but directionally consistent with the meta-analytic crossover). Also confirmed within the family: the worked-example effect (novices learn more studying a worked solution than solving cold; Sweller & Cooper 1985; g = 0.48 in Barbieri et al. 2023).
Design consequence. Five rules, each traceable to a number above:
- The content decides, then the learner dials (Willingham's rule made data — see the viz hint below): a manipulable is built when the node's causal structure rewards one, never to please a "style."
- Never a bare sandbox: every manipulable ships inside a predict → act → explain micro-cycle (self-explanation g = 0.46), with scaffolds in the artifact (g+ = 0.60), prediction gates as the guidance rail (guided-discovery cell, d = +0.30).
- Interactivity is spent on the model, not the chrome (control per se g = 0.05): the learner advances between segments; within a segment the dynamic runs itself; no scrub-theater.
- No text over motion; no decoration ever (redundancy; seductive details g ≈ −0.16…−0.33): explanation text sits before or after the dynamic, and every pixel either teaches or is deleted.
- Scaffolds fade with measured knowledge (expertise reversal, +0.505/−0.428): novice-state nodes open with a worked drive of the model (watch one demonstration run under a "what happens next?" gate) before free manipulation unlocks; comfortable-state nodes go straight to open manipulation. Engram already tracks the exact signal needed to titrate this — the node's own state, pretest result, and lapse count.
What the audit killed (do not build on these)
- "Statics are safest for factual retention; interactivity pays off mainly for transfer." Refuted 0-3 by the verifiers against the cited corpus. Media-by-outcome routing of that kind stays out of Engram.
- "Expertise reversal is weaker than believed / merely ordinal." Refuted (the 2025 meta shows a genuine disordinal crossover with both subgroup CIs excluding zero). What is true: Kalyuga's 2007 narrative mid-range (1.72) roughly doubles the meta-analytic difference (0.971) — narrative-review inflation is real; the effect is not.
What remains honestly open (and what v0.5 does about each)
Nothing in these four areas survived adversarial verification — which means no verifiable evidence was found, not that the answer is no. Each gets a conservative design stance:
-
Visual encoding × verbal retrieval (does dual-coded encoding survive verbal free-recall testing? do diagram-completion/sketch-recall formats work as retrieval practice?). → Reviews stay exactly as they are: verbal free recall, graded against the rubric. No visual review formats ship until evidence does. The explorable's embedded retrievals are prose productions for precisely this reason.
-
n-of-1 medium measurement (is comparing one learner's retention across encoding media methodologically sound?). →
stats.modalityships as suggestive personal telemetry, guarded by the same per-arm floor as n-of-1 experiments (≥6 first-reviews per arm), narrated with its n, and never called proof. It informs the dial; it never overrides the content rule.The confound, stated plainly (found in a v0.5.0 dogfood session, documented in v0.5.1). The arms are not randomized, and cannot be without violating the content rule this very document establishes. Explorables are routed to threshold and high-affordance concepts on purpose; the dialogue arm therefore fills with the remaining material. Under
threshold-onlythe explorable arm is exactly the topic's portal concepts — the hardest nodes in the graph. Undereagerit is every node whose content is visually affordant. Either way the comparison carries medium and material together, and a lower explorable-arm recall may simply mean explorables were spent on the hard things. So: the engine ships the caveat inside the stats block (modality.caveat), the dashboard prints it beside the bars, and/coachmust voice it whenever it reports the number. What the telemetry can honestly detect is a large, stable divergence — never a small one, and never a causal claim. A properly randomized answer would need theexperimentmachinery (assign comparable nodes to media at random within one affordance class); that is future work, and it is the honest form of the question. -
Preference-as-engagement value (does honoring a visual preference buy consistency even without a learning-rate edge?). → Preference is honored as autonomy (the
visualsdial, the ask-once offer, on-request builds) under P10's existing logic — consistency dominates — without claiming a retention mechanism the evidence doesn't show. -
LLM-generated artifact efficacy + mnemonic-medium field data (2023–2026 frontier, Quantum Country). → Engram's own receipts are the instrument: registration + medium-stamped receipts make every install a small honest field study of exactly this question. The Phase-2 exit criterion (
docs/04) finally has its instrument.
(Also still open from the general corpus: drawing-to-learn, predict-observe-explain, and gesture beyond the Mayer-corpus self-explanation g = 0.46 and embodiment g = 0.35 entries. The Contract keeps prediction and self-explanation — which are verified — and does not add sketch-input widgets on vibes.)
The machinery (v0.5) — each piece traced to its principle
| Piece | What it is | Traces to |
|---|---|---|
viz hint on nodes | The curriculum architect tags every node with content-declared visual affordance: {affordance: high|some|none, kind: dynamic-process|causal-parameter|structural|distributional|procedural|comparative, hook: "<the manipulation that would kill the likely misconception>"}. The engine stores it opaquely; skills own semantics. | Willingham's rule (01 §Rejections); representational-vs-decorational d = 0.40 / −0.05; procedural d = 1.06 |
visuals dial | off · threshold-only (default, byte-compatible with v0.4) · eager (threshold nodes and affordance: high nodes). Any level: the learner can request an explorable on any node, any time — autonomy override, same shape as "just tell me." | P10 autonomy; preference-as-engagement (open Q3, conservative) |
| Ask-once offer | Under threshold-only, the first time a topic hits a high-affordance non-threshold node, the tutor may offer once (arrow-key): build this one / always (visuals eager) / no — then silence for the rest of the topic. Consent rule: every dial change is offered with its evidence, applied only on yes. | Constitution art. 9 (open model); calm-surface doctrine |
| Worked-drive scaffold gate | Contract v2: novice-state nodes open the manipulable with one demonstrated run under a prediction gate before free manipulation unlocks; comfortable-state nodes manipulate immediately. Scaffold level comes from node state + pretest + lapses. | Expertise reversal +0.505/−0.428; worked examples g = 0.48; "provide when in doubt" |
| Artifact registration | engram.py artifact set/clear/list — the graph's artifact field is engine-owned (validated file, home-relative path, survives --replace, payload-supplied values stripped). The smith registers what it builds; doctor notes unregistered files and fails dangling registrations. | Article 10 (receipts or it didn't happen); Contract clause 7 |
| Medium-stamped receipts | Every rate/receipt stamps whether the node had a registered explorable at grading time — evidence of the medium can never be rewritten retroactively. | Article 10 |
stats.modality | First-review recall, explorable-encoded vs dialogue-only, one datum per node, ≥6 per arm before any verdict; read is ahead/behind/indistinguishable/insufficient-data, and every read ships caveat (the arms are not randomized — see the confound above). Narrated by /coach with its n and its caveat; shown on the dashboard; never a proof, always the learner's own numbers. | Article 7 (adapt on evidence, never taxonomy); open Q2, conservative |
The invariants, so this can be checked against them (same discipline as docs/05):
- The engine's pedagogy is untouched. FSRS math, receipts-before-state, the blind assessor, free-recall probes — identical. v0.5 adds registration, stamping, and read-only telemetry.
- Defaults are byte-compatible. A v0.4 learner model self-heals to
artifacts: threshold-onlyand behaves exactly as before; every new behavior is opt-in (eager), on-request, or invisible plumbing. - The content rule outranks the preference rule.
visuals eagerstill builds only for high-affordance or threshold nodes — there is no setting that decorates anaffordance: nonenode. - Retention is still the north star. The modality telemetry exists so the medium can lose: if a learner's explorable-encoded nodes hold worse, the coach says so with the numbers and offers the dial down. The founder's beloved medium submits to the founder's own constitution — and the instrument that judges it declares its own confound rather than flattering the feature that built it.
The founding question, answered
Q: "I'm a huge visual learner — interactive HTML made features click for me in seconds. But this might be just me. Should Engram lean into it?"
A: The click was real, the label is wrong, and the lean-in is earned — under discipline. What happened with the draggable face was not a "visual learner" being served their style (styles-matching remains dead — 01 §Rejections stands): it was a high-affordance concept finally meeting its content-matched medium — "features as manipulable dimensions" is exactly the causal-parameter structure where guided manipulables carry their largest verified effects — wrapped, crucially, in the thing the evidence says does the work: you acted on it and watched the consequence. That mechanism is universal cognitive architecture, which is why Engram now lets the content declare the affordance (the architect's viz hint) and any learner dial the eagerness. Your preference for the medium is honored as autonomy and motivation; your retention receipts — not your enthusiasm, and not this document — get the final word on whether it earns its keep for you. That is the same bargain as everything else in Engram: derive what can be derived, memorize only the arbitrary, test everything, schedule everything — and now: make manipulable what is truly manipulable, guide every manipulation, and let your own data arbitrate.