README.md

September 8, 2025 · View on GitHub

Supported functions

Real-time Speech recognitionSpeech synthesisVoice activity detection
✔️✔️✔️

Supported platforms

ArchitectureAndroidiOSWindowsmacOSlinux
x64✔️✔️✔️✔️
x86✔️✔️
arm64✔️✔️✔️✔️✔️
arm32✔️✔️
riscv64✔️

Supported programming languages

1. C++2. C3. Python4. JavaScript
✔️✔️✔️✔️
5. Go6. C#7. Kotlin8. Swift
✔️✔️✔️✔️

It also supports WebAssembly.

Introduction

This repository supports running the following functions locally

  • Streaming speech-to-text (i.e., real-time speech recognition)
  • Text to speech (e.g., vits models from piper)
  • VAD (e.g., silero-vad)

on the following platforms and operating systems:

with the following APIs

  • C++, C, Python, Go, C#
  • Kotlin
  • JavaScript
  • Swift

We support all platforms that ncnn supports.

Everything can be compiled from source with static link. The generated executable depends only on system libraries.

HINT: It does not depend on PyTorch or any other inference frameworks other than ncnn.

Please see the documentation https://k2-fsa.github.io/sherpa/ncnn/index.html for installation and usages, e.g.,

  • How to build an Android app
  • How to download and use pre-trained models

We provide a few YouTube videos for demonstration about real-time speech recognition with sherpa-ncnn using a microphone:

DescriptionURL
Streaming speech recognitionAddress

https://github.com/k2-fsa/sherpa-ncnn/releases/tag/models

How to reach us

Please see https://k2-fsa.github.io/sherpa/social-groups.html for 新一代 Kaldi 微信交流群 and QQ 交流群.

See also