Testing Guide

July 29, 2026 · View on GitHub

SCPN Phase Orchestrator builds its Python test surface around dedicated module-owned tests. Measured coverage is 94.34% line / 93.22% branch (CI lanes; see the V&V Report §1.1); the authoritative gate is a per-domain no-decrease ratchet enforced by tools/coverage_guard.py, not a flat percentage. The 60% figure in older notes was the floor held during the dedicated-test-surface rebuild and is superseded. Each production module regains coverage through its own focused unit, property, parity, or pipeline tests.

Running Tests

# Full suite
py -3.12 -m pytest tests/ -v --tb=short

# Single file
py -3.12 -m pytest tests/test_prop_lyapunov_dimension.py -v

# Only property-based tests
py -3.12 -m pytest tests/test_prop_*.py -v

# With coverage
py -3.12 -m pytest tests/ --cov=scpn_phase_orchestrator --cov-report=term-missing

What this test surface is designed to prove

This test surface is organized around decision risk, not only pass/fail:

  • Numerical risk is covered by property checks and degenerate edge fixtures.
  • Cross-engine consistency risk is covered by parity suites and analytical limits.
  • Runtime risk is covered by stress profiles and bounded slow tests.
  • Regression risk is covered by mutation testing and module-level ownership checks.

For a production-affecting change, treat these four checks as minimum evidence:

  1. property and edge tests for the changed module family,
  2. parity or analytical checks for any coupled path,
  3. at least one replay or slow validation path on realistic scale,
  4. release-note and docs updates for changed behavior.

This ordering mirrors the internal release checklist: mathematics and safety signals are validated first, throughput and scale are validated next, and documentation is updated before any external claim changes.

Hypothesis Profiles

The project defines hypothesis profiles in pyproject.toml:

Profilemax_examplesUse case
dev50Local development (default)
ci500CI pipeline, thorough

Select a profile:

py -3.12 -m pytest tests/ --hypothesis-profile=ci

Test Architecture

Property-Based Tests (test_prop_*.py)

These are computational theorem provers. Each @given test generates 50-500 random inputs and verifies that a mathematical invariant holds for all of them. If any counterexample is found, hypothesis shrinks it to the minimal failing case.

FileTestsWhat it proves
test_prop_lyapunov_dimension.py44Lyapunov spectrum: length=N, sorted descending, finite. Kaplan-Yorke D_KY ∈ [0,N]. Correlation integral monotonic in ε.
test_prop_basin_stability.py28S_B ∈ [0,1], n_converged ≤ n_samples. Multi-basin threshold monotonicity. Strong coupling → high S_B.
test_prop_entropy_transfer.py25EPR ≥ 0, TE ≥ 0, TE diagonal = 0. TE-adaptive coupling preserves zero diagonal and non-negativity.
test_prop_hodge_spectral.py30Laplacian: PSD, row sums = 0. Fiedler λ₂ > 0 iff connected. Hodge: gradient + curl + harmonic = total.
test_prop_recurrence_rqa.py26Recurrence matrix: symmetric, diagonal = True. RR, DET, LAM ∈ [0,1]. Cross-recurrence shape and bounds.
test_prop_chimera_winding.py19Chimera index ∈ [0,1], coherent/incoherent disjoint. Winding numbers integer-valued, reverse ≈ negation.
test_prop_free_energy_boltzmann.py25Boltzmann weight ∈ (0,1] for U ≥ 0, monotonic in U and T. SSGF costs: c1 ∈ [0,1], c3 ≥ 0, c4 = 0 for symmetric W.
test_prop_embedding_poincare.py15Delay embedding shape = (T-(m-1)τ, m). Optimal delay ≥ 1. Optimal dimension ∈ [1, max_dim].
test_prop_ei_balance_npe.py18Phase distance: symmetric, diagonal = 0, values ∈ [0,π]. NPE ∈ [0,1], sync → 0. EI ratio ≥ 0.
test_prop_simplicial_reduction.py8σ₂ = 0 reduces to standard Kuramoto (exact match). σ₂ ≠ 0 differs.
test_prop_swarmalator_inertial.py10Swarmalator: J=0 decouples phase from position. Inertial: θ wrapped to [0,2π).
test_prop_plasticity_stochastic.py22Eligibility: symmetric, ∈ [-1,1], zero diagonal. StochasticInjector: D=0 → no change, output ∈ [0,2π).

Degenerate Edge Cases (test_degenerate_edges.py)

98 tests that push every engine to its boundaries: N=1 oscillator, dt=0, zero coupling (free rotation), identical phases, phase wrapping at 0 and 2π, extreme coupling strengths, negative frequencies. Parametrised across all 5 engine types (UPDE, Stuart-Landau, Simplicial, Swarmalator, Inertial).

Cross-Module Roundtrips (test_roundtrip_consistency.py)

86 tests that verify mathematical consistency across module boundaries:

  • Synchronised phases → R ≈ 1, PLV ≈ 1, NPE ≈ 0, chimera_index ≈ 0 (four independent measures agree)
  • Spectral λ₂ predicts synchronisability → verified by simulation
  • Projection roundtrip: project_knm always produces valid K_nm
  • Simplicial σ₂ = 0 roundtrip: reduces to standard Kuramoto exactly
  • Free rotation → analytical winding number matches
  • Transfer entropy: directional, correct shape
  • NPE vs R anti-correlation across synchronisation spectrum

Module Tests

Dedicated test files for each subsystem covering unit-level behaviour, input validation, edge cases, and dataclass contracts:

SubsystemFilesModules tested
SSGFtest_ssgf_modules.pyGeometryCarrier, CyberneticClosure, EthicalCost
UPDE mathtest_upde_math.pyTorusEngine, IntegrationConfig, check_stability, OttAntonsenReduction
Couplingtest_coupling_modules.pyLagModel, UniversalPrior, KnmTemplateSet
Driverstest_drivers_oscillators.pyPhysicalDriver, PhaseQualityScorer, CoherenceMonitor
Supervisortest_supervisor_modules.pyEventBus, RegimeManager, InformationalDriver, SymbolicDriver
Imprinttest_imprint_actuation.pyImprintModel, ActionProjector, ActuationMapper
Bifurcationtest_bifurcation.pytrace_sync_transition, find_critical_coupling

Writing New Tests

Property test pattern

from hypothesis import given, settings
from hypothesis import strategies as st

class TestMyInvariant:
    @given(
        n=st.integers(min_value=2, max_value=12),
        seed=st.integers(min_value=0, max_value=200),
    )
    @settings(max_examples=50, deadline=None)
    def test_output_bounded(self, n: int, seed: int) -> None:
        rng = np.random.default_rng(seed)
        phases = rng.uniform(0, TWO_PI, n)
        result = my_function(phases)
        assert 0.0 <= result <= 1.0

Key conventions:

  • deadline=None for tests that run simulations (Lyapunov, basin stability)
  • Small N (2-16) for CPU speed — property tests run 50-500 iterations
  • suppress_health_check=[HealthCheck.too_slow] for Monte Carlo tests
  • Use _connected_knm(n, seed=seed) helper for reproducible symmetric coupling matrices
  • Tolerances for float comparison: atol=1e-12 for exact, atol=1e-10 for simulation

Degenerate edge test pattern

@pytest.mark.parametrize("n", [2, 4, 8])
def test_zero_coupling_free_rotation(self, n: int) -> None:
    eng = UPDEEngine(n, dt=0.01)
    # ... verify analytical prediction under extreme conditions

Cross-Engine Parity Tests (test_engine_parity.py)

The parity matrix verifies that engines which should agree on a given scenario actually produce the same result:

Engine AEngine BScenarioTolerance
UPDE EulerTorusEngineSingle step, small dt1e-4
UPDE EulerSplittingEngineSingle step, small dt1e-3
UPDE EulerRK4500-step converged R0.05
Simplicial σ₂=0UPDE EulerAny input (hypothesis)1e-10
UPDE / Torus / SplittingAnalyticalFree rotation θ = ωt1e-6

Plus analytical validation:

  • Spectral K_c: K > 2K_c → sync, K < K_c/10 → no sync
  • Stuart-Landau: r → √μ (property-based, μ ∈ [0.1, 5.0])
  • OA vs UPDE: Lorentzian g(ω), above and below K_c

Stress / Scale Tests (test_stress_scale.py)

Production-scale validation marked with @pytest.mark.slow:

TestScaleVerifies
N=1000 identical sync1000 oscillatorsR > 0.90 after 1000 steps
N=1000 random R1000 oscillatorsR < 0.15 (≈ 1/√N)
N=1000 NPE1000 oscillatorsFinite, no OOM
N=512 Laplacian512×512 matrixPSD, Fiedler > 0
10k steps no drift16 osc, 10000 stepsAll finite, R ∈ [0,1]
50k steps stable8 osc, 50000 stepsR variance < 0.1

Run slow tests explicitly: py -3.12 -m pytest -m slow

Engine Rigor Tests (test_engine_rigor.py)

Dedicated validation for auxiliary engines:

EngineTestsKey invariants
HypergraphEngine5k-body coupling, free rotation, output bounds
Market module5Hilbert phase extraction, R shape, regime detection
Envelope solver6Shape, non-negative, modulation depth ∈ [0,1]
Adjoint gradient4cost_R bounds, gradient shape, zero diagonal
DelayBuffer/Engine7Push/get, early access, delay=1 ≈ standard

CI Integration

CI runs the main suite on Python 3.11–3.13 and collects main-line coverage on Python 3.12. Separate FFI jobs exercise supported interpreter/OS combinations. The Python fallback uses pure-NumPy integrators; the Rust path uses spo-kernel via PyO3. Tests handle both paths — see test_degenerate_edges.py::TestUPDEZeroDt for the pattern.

Coverage gate: a per-domain no-decrease ratchet (tools/coverage_guard.py) against the CI coverage lanes, seeded from the measured baselines in tools/coverage_guard_thresholds.json (line, ≥93% global) and tools/coverage_guard_branch_thresholds.json (branch, ≥91% global). The floors ratchet upward from each green run and never decrease; new modules ship at 100% and lift their domain's floor. (The historical "60% minimum" was the rebuild floor, now superseded.)

Convergence & Topology Tests (test_convergence_topology.py)

Numerical and graph-theoretic proofs:

  • Convergence order: Euler and RK4 exact on free rotation (linear ODE); coupled case: RK4 more accurate than Euler at same dt
  • Topology dynamics: all-to-all fastest sync; star hub entrains spokes; ring λ₂ > chain λ₂ (algebraic connectivity proof); disconnected → no sync
  • Delay τ→0 limit: delay_steps=1 converges like standard UPDE; large delay (50 steps) destabilises sync
  • Benchmark baseline: 1000 steps at N=32 in <5s; order parameter <1ms at N=256

Mutation Testing (test_mutation_killers.py)

Mutation testing injects small bugs (mutants) into source code and checks whether the test suite catches them. A survived mutant means the tests have a blind spot. We use mutmut v2.4.5 running on Kaggle (Linux kernel) since mutmut does not support Windows natively.

Results (2026-03-28)

ModuleMutantsSurvivedKilled by new tests
upde/order_params.py281622 killer tests
upde/numerics.py10510 killer tests

All 21 real survivors are now covered by dedicated tests in test_mutation_killers.py. The tests target specific operator and value mutations that the existing suite missed:

  • Boundary returns: phases.size == 0(0.0, 0.0) (exact zeros, not just "small")
  • Imaginary unit: exp(1j * theta) — verify 1j not mutated to 1
  • Operator semantics: max_omega + max_coupling (sum, not max)
  • Exact defaults: every IntegrationConfig default value asserted exactly
  • PLV edge cases: empty arrays, size mismatch, anti-phase locking

Running mutation tests

mutmut requires Linux. On Kaggle or WSL:

# Install mutmut v2 (v3 changed the CLI)
pip install mutmut==2.4.5

# Run on a single module with targeted fast tests
mutmut run \
  --paths-to-mutate src/scpn_phase_orchestrator/upde/order_params.py \
  --tests-dir tests/ \
  --runner "python -m pytest tests/test_mutation_killers.py -x -q --tb=no" \
  --no-progress

# Show survivors
mutmut results

The Kaggle kernel anulum/spo-mutmut-v2 is configured for batch mutation testing across multiple modules.

Release governance from test surfaces

For each release train, this section is treated as a pre-merge control list:

  • confirm module-owned unit and property tests for all modified production modules,
  • confirm parity tests for every execution path in scope,
  • confirm slow/scale suites were run for benchmark or performance-affecting edits,
  • confirm mutation coverage updates were included when risk-sensitive code changed.

The result is not just a pass/fail signal. It is evidence that a regression in control, safety, or replay logic is likely to be detected before deployment.