AudioCPP Command Usage
August 10, 2026 ยท View on GitHub
Use audiocpp_cli for direct model inference.
audiocpp_cli --task <task> --family <family> --model <model-dir> --backend <backend> [inputs] [outputs]
Common Options
| Option | Values | Default | Meaning |
|---|---|---|---|
--task | gen, tts, clon, vc, svc, s2s, asr, align, vad, diar, sep, vdes, midi | required | User task. |
--family | model family name | required | Selects the model implementation. Must match a registered loader (audiocpp_cli --list-loaders). |
--model | local model directory | required | Path to local model assets. |
--backend | cpu, cuda, vulkan, metal, best | cpu | Inference backend. |
--mode | offline, streaming | offline | Run mode. Most models are offline. |
--device | integer | 0 | Backend device index. |
--list-devices | flag | off | List available backend devices and exit; combine with --backend/--device to select one. |
--threads | integer | 4 | Backend/OpenMP worker threads. |
--log | flag | off | Print progress and timing logs to stdout. |
--log-file | path | not set | Stream progress and timing logs to a file. |
--metrics | flag | off | Print compact offline wall time, audio duration, RTF, realtime speed, sample rate, and channel metrics. |
Common Inputs And Outputs
| Option | Used by | Meaning |
|---|---|---|
--text | generation, TTS, ASR context, alignment transcript | Input text. |
--audio | generation/editing, ASR, VAD, diarization, separation, conversion, alignment | Input WAV, or - to stream raw PCM from stdin (requires --mode streaming). |
--input-format | streaming ASR with --audio - | Raw PCM sample format, s16le or f32le. Default s16le. |
--input-rate | streaming ASR with --audio - | Raw PCM sample rate in Hz. Default 16000. |
--input-channels | streaming ASR with --audio - | Raw PCM channel count. Default 1. |
--voice-ref | voice clone / voice design / some VC paths | Reference voice WAV. |
--language | language-aware models | Language code. |
--out | single-primary-output models | Output file path, such as WAV for audio tasks or MIDI/JSON for MuScriptor. |
--out-dir | multi-output or batch models | Output directory. |
--segments-out | VAD | Speech segments JSON. |
--vad-chunks-out | offline VAD | VAD-based audio chunk windows JSON. |
--turns-out | diarization | Speaker turns JSON. |
--words-out | ASR/alignment | Word timestamps JSON. |
--audio-chunk-seconds | ASR | Split long audio before model inference, where supported. |
--audio-chunk-mode | ASR/alignment | auto, fixed, vad, or none, where supported. |
Common Generation Options
Omit these unless you need explicit control. If --seed is omitted, models that sample use a random seed.
| Option | Values | Meaning |
|---|---|---|
--seed | integer | Reproducible random seed. |
--max-tokens | integer | Maximum generated tokens for AR/LLM-style models. |
--max-steps | integer | Maximum diffusion or generation steps for models that expose it. |
--temperature | float | Sampling temperature. |
--top-k | integer | Top-k sampling limit. |
--top-p | float in (0, 1] | Nucleus sampling limit. |
--repetition-penalty | float | Penalize repeated tokens. |
--do-sample | true, false | Enable sampling instead of greedy decode. |
--guidance-scale | float | Classifier-free guidance scale. |
--num-inference-steps | integer | Diffusion/flow denoising steps. |
--text-chunk-size | integer chars | Split long text where supported, including TTS text. |
--text-chunk-mode | default, tag_aware, japanese, endline | Select text chunking mode where supported. |
Batch Inputs
| Option | Meaning |
|---|---|
--request-sequence <json> | Run multiple JSON requests through one offline model session. |
--batch-text-file <txt> | One request per non-empty text line. |
--batch-text-dir <dir> | One request per .txt, .md, or .json file; each file is normalized into a single paragraph. |
--batch-audio-dir <dir> | One request per .wav file. |
--batch-audio-role audio|voice_ref|source_audio|target_voice|prosody_ref|style_ref | How to use each batch WAV. |
--batch-merge-audio none|concat | Keep outputs separate or concatenate generated audio. |
--batch-manifest-out <json> | Write a batch output manifest. |
--batch-text-dir reads .txt and .md files as plain text. For .json, use either a JSON string root or an object with a string input or text field.
Use --request-sequence when you want to send multiple requests in one long-lived offline session:
audiocpp_cli --task tts --family pocket_tts \
--model models/PocketTTS-GGUF/english/pocket-tts-english-q8_0.gguf \
--backend cuda \
--request-sequence requests.json \
--out-dir outputs \
--metrics
For each request id, --metrics prints metrics[<id>].wall_ms, audio_duration_ms, rtf, x_realtime, sample_rate, and channels.
Model Docs
| Need | Doc |
|---|---|
| Speech, voice clone, long-form TTS | tts.md |
| Music and sound generation | music_generation.md |
| OmniVoice TTS, voice cloning, voice design, and streaming | models/omnivoice.md |
| ASR models | asr.md |
| VAD and diarization | speech_analysis.md |
| Audio tools, voice conversion, codec, and source separation | audio_tools.md |