Roadmap
June 24, 2026 · View on GitHub
Where this project is headed. The aim is to grow from a collection of helper skills into a Materials Simulation Skill Protocol — a small, opinionated standard layered on the open Agent Skills spec that makes computational-materials and numerical-simulation skills portable, composable, and scientifically trustworthy across any compatible agent.
The guiding bet: validated, portable skills beat one monolithic simulation container.
This file is the general direction. Day-to-day bug fixes, doc corrections, and small improvements are ordinary contribution work — see CONTRIBUTING.md — not roadmap items.
Guiding principles
- Start small, prove it. One well-tested script beats a broad, fragile skill.
- Progressive disclosure. Detail lives in
references/, notSKILL.md(keepSKILL.mdunder ~500 lines). - Every claim is checkable. Numeric examples and documented output fields are
guarded by deterministic
script_checks; rules cite an authoritative standard. - Honesty over polish. A Security section documents only what the code does; a limitation is stated, not hidden.
- Composability first. Skills reference and build on each other rather than duplicate logic.
The protocol (conventions on top of the Agent Skills spec)
These conventions are what turn a collection into a coherent standard. They are
codified in docs/PROTOCOL.md with a conformance checklist;
the frontmatter, naming, Security-subsection, and eval-coverage rules are enforced
in CI by mss validate and mss eval. The direction is to enforce more of the
checklist automatically (e.g. the script I/O envelope and compatibility).
- Script I/O envelope. Every CLI emits
--jsonas{ "inputs": {...}, "results": {...}, "notes": [ ... ] }, with exit2on bad input,1on runtime error,0on success. - Units & provenance. Numeric results carry units; outputs that feed FAIR packaging expose engine/version/standard references.
- Cited science. Each skill names the standard or reference behind its rules (e.g., ASME V&V20, SchedMD, CMSO IRIs) so an agent can defend an answer.
- Declared compatibility. Use the spec's
compatibilityfrontmatter field for environment needs (Python version, NumPy/SciPy, network) instead of prose. - Spec-conformant metadata. Keep
metadataspec-aligned and runskills-ref validatefor conformance alongside the repo's own validator. - Security tiers that match reality. A skill either declares the
Bashit needs to run its scripts, or it ships logic the agent applies without execution — never both stories at once. Standard## Securitysubsections. - Composition contract. A skill may declare sibling skills it builds on
(a lightweight
metadata.depends_on), enabling a dependency graph and bundles.
Direction
Evaluation as a first-class gate
The three-layer evaluation harness exists: a deterministic
script_checks CI gate plus the agent-agnostic skill-evaluator skill (trigger
eval + with/without quality benchmarking across Claude Code, Codex, Antigravity,
Cursor, Copilot, Amp, opencode, Grok). The direction now is depth, not foundation:
- Grow deterministic
script_checkstoward every eval case with a computable answer, so doc↔code drift can never regress silently. - Wire the LLM-judge (Layer 3) into CI as an opt-in gate — when a CLI and key are available, gate skill changes on "with-skill beats baseline."
- Precise per-CLI trigger detection (parse tool-use events) beyond the current cross-tool heuristic; per-skill dashboards tracking pass-rate and token/latency deltas across iterations so regressions and bloat are visible.
- Cross-skill integration scenarios: graded end-to-end "campaigns" exercising several skills together (stability → mesh → solver → convergence → validate → package).
Breadth — new skill areas
Breadth is gated by the verifiable-correctness principle: we expand first where correctness is checkable without an expert, and route expert-judgment domains through expert contribution. Each new skill starts from one well-tested script.
- Self-validating first (checkable against math/standards — the safest growth): richer manufactured-solution and benchmark libraries with known exact answers, spectral / multigrid methods, uncertainty quantification on analytical test functions, deeper V&V, more scheduler portability (PBS/LSF).
- Expert-reviewed domains (correctness rests on judgment — only via expert contribution + review): DFT input review (VASP/QE — k-points, cutoffs, smearing), LAMMPS/MD input review, interatomic-potential / MLIP-readiness selection, CALPHAD/thermodynamic sanity.
- Data / FAIR: tighter NOMAD / OPTIMADE / Materials Project round-tripping (verifiable against the published APIs/schemas).
Interoperability & distribution
- Skill registry / index. A generated, machine-readable index (name, description, category, compatibility, dependencies, cited standards) so agents and humans can discover and compose skills.
- Skill bundles. Named sets (e.g.,
phase-field-starter,dft-campaign) installable as a coherent group viamss install. - Versioning & governance. Semver per skill, enforced CHANGELOG discipline, and a protocol-conformance badge for skills that pass the checklist.
- Publishing. Package conformant skills for the agentskills ecosystem so they
install cleanly into Claude Code, Codex, Antigravity (
agy, the CLI that replaced Gemini CLI on 2026-06-18), Cursor, Copilot, Amp, opencode, and Grok.
Contribution principles
- Start small: one well-tested script beats a broad but fragile skill.
- Preserve progressive disclosure: detailed tables go in
references/. - Every numeric claim in a
SKILL.mdgets ascript_check; every rule cites a standard. - Add eval cases with concrete prompts and assertions for every skill, and a
deterministic
script_checkwhenever the answer is computable.