NepTrain

August 21, 2026 · View on GitHub

English | 简体中文

NepTrain

NepTrain is a command-line tool for the complete lifecycle of neuroevolution potential (NEP) models. It can run training, molecular dynamics, structure selection, and labeling as standalone tasks, or compose the same steps into a resumable active-learning workflow.

train → validate → explore → select → label → evaluate → update

Standalone commands and automated workflows use the same scientific adapters and execution targets. The workflow layer owns planning, state transitions, and acceptance; it does not duplicate the scientific calculation logic.

What NepTrain provides

  • GPUMD or TorchNEP training with a canonical nep.txt result.
  • GPUMD or LAMMPS sampling, including LAMMPS DynSpin for spin MD.
  • VASP, ABACUS, MACE, DeepMD/DPA, or TACE labeling.
  • Element-set-aware farthest-point sampling with deterministic provenance.
  • Local process, local Slurm, and SSH + Slurm execution targets.
  • Content-addressed task bundles, atomic publication, hash validation, and resumable workflow state.
  • PNG training-convergence and evaluation reports generated with Matplotlib.

Installation

NepTrain requires Python 3.10 or later:

pip install NepTrain

Install only the optional runtime you need:

# TorchNEP training
pip install torch
pip install 'NepTrain[torchnep]'

# MACE teacher labeling
pip install 'NepTrain[mace]'

# DeepMD / DPA teacher labeling
pip install 'NepTrain[deepmd]'

# TACE teacher labeling (TACE is currently installed from its source repository)
pip install \
  'TACE @ git+https://github.com/xvzemin/tace.git@4b977dcc13ee87d8ba6cceba3ffb7abe43c087c8'

# Optional cuEquivariance acceleration on Ampere-or-newer CUDA 12 GPUs
pip install \
  'TACE[cueq12] @ git+https://github.com/xvzemin/tace.git@4b977dcc13ee87d8ba6cceba3ffb7abe43c087c8'
export TACE_USE_CUE=1

# SOAP descriptors when manual selection has no NEP model
pip install 'NepTrain[soap]'

LAMMPS, GPUMD, VASP, ABACUS, and first-principles resource files are supplied by the user or computing platform.

First successful run

Use the deterministic toy workflow to check the NepTrain installation without submitting DFT or Slurm jobs:

neptrain smoke --profile ordinary

For a real project, validate every configured execution target before submission:

neptrain doctor --project project.yaml

doctor reads the project backends, stage targets, setup scripts, and target environment. It checks the actual GPUMD/TorchNEP, LAMMPS/GPUMD, VASP/ABACUS, or MACE/DeepMD/TACE runtime required by each target.

Standalone steps

Train one model:

neptrain train train.xyz \
  --backend torchnep \
  --config nep.in \
  --device cuda \
  -o nep.txt

Run a structure × temperature MD batch:

neptrain md structures/ \
  --backend lammps \
  --model nep.txt \
  --temperature 300 500 700 \
  --steps 100000 \
  --max-concurrent 12 \
  -o trajectories.xyz

Select representative structures:

neptrain select md-300.xyz md-600.xyz \
  --base train.xyz \
  --nep nep.txt \
  --max-selected 64 \
  --out selected.xyz \
  --report selected.selection.json

Label structures with VASP:

neptrain label candidates.xyz \
  --backend vasp \
  --input-file INCAR \
  --resources /shared/potpaw_PBE \
  --potcar-manifest vasp-resources.json \
  --structures-per-job 1 \
  --max-concurrent 20 \
  -o labeled.xyz

Submitted standalone tasks use one control surface:

neptrain task status runs/label-...
neptrain task logs runs/label-...
neptrain task wait runs/label-...
neptrain task retry runs/label-...
neptrain task cancel runs/label-...

The final output is published only after every shard passes identity, order, and artifact validation. Existing output is not overwritten unless --force is given.

Automated workflow

Create a schema-v8 project:

neptrain workflow init \
  --profile slurm \
  --ensemble npt \
  --dft-backend vasp \
  --directory fe-project
cd fe-project

After filling in structures, training data, templates, resource manifests, and execution targets:

neptrain doctor --project project.yaml
neptrain workflow run project.yaml --prepare-only
neptrain workflow run fe-workflow

Inspect and control it with:

neptrain workflow status fe-workflow --jobs
neptrain workflow resume fe-workflow
neptrain workflow restart fe-workflow --generation 3 --from label --dry-run
neptrain workflow stop fe-workflow
neptrain workflow extend fe-workflow 5

Human-readable workflow status focuses on the active generation, a compact temperature path with observable MD progress in ps, and a per-generation RMSE table. --jobs groups large MD and labeling batches by generation, stage, and attempt; --json retains every individual execution record.

NepTrain accepts only schema_version: 8. Unknown fields and legacy project formats fail explicitly instead of being migrated silently.

Built-in LAMMPS and GPUMD sampling adapts trajectory dump spacing to the run length. Custom LAMMPS templates can use {{ dump_interval }}. Slurm analysis targets can also declare memory_ladder: [4G, 8G, 16G]; selection OOM advances to the next tier with a bounded, traceable retry.

Optional Feishu progress notifications can be configured directly in project.yaml:

notifications:
  feishu:
    webhook: https://open.feishu.cn/open-apis/bot/v2/hook/REPLACE
    secret: REPLACE
    timeout_seconds: 5

neptrain doctor --project project.yaml sends a real connectivity probe. Workflow delivery runs on a background thread and is always best effort: notification failures never change workflow state, exit status, or the scientific ledger. Progress and terminal messages include both the workflow id and absolute workflow path so concurrent runs remain distinguishable.

Labeling backends

BackendRuntimeMain boundary
VASPUser-provided VASP and pinned POTCAR manifestOrdinary energy/force/virial labels
ABACUSUser-provided ABACUS and pinned UPF/ORB manifestOrdinary or DeltaSpin spin/mforce labels
MACENepTrain[mace]Ordinary structures; no mforce
DeepMD / DPANepTrain[deepmd]DPA-3 and supported DPA-4 formats; no mforce
TACEOfficial TACE source installCheckpoints supported by tace-eval; spin requires real noncollinear magnetic-force output

Teacher-model labeling still enters through neptrain label. The hidden model-worker command is an internal runner protocol, not a second user-facing CLI.

Documentation and tutorials

The website documentation is built from one Chinese source tree with complete English translation catalogs. Both languages are checked with strict Sphinx warnings in CI.

Scientific-data boundaries

  • Structure identity uses the versioned structure-id.v3 contract.
  • Spin datasets use canonical spin and mforce extxyz arrays.
  • VASP and ABACUS resources are pinned by relative path and SHA256 manifest.
  • Model, dataset, task, result, and publication identities are content addressed.
  • Workflow ledger entries are scientific commit points; controller state is execution intent only.

Support