README.md

August 19, 2026 ยท View on GitHub

Kokoro TTS Header


A Home Assistant custom integration for connecting to Kokoro FastAPI, enabling high-quality local Text-to-Speech. Easily send TTS audio to your speakers or media players directly from Home Assistant.

๐ŸŽง Listen to a preview: โ–ถ Play



โœจ Features

  • ๐Ÿ”Š Convert text to speech using Kokoro FastAPI
  • โšก Low-latency responses for near real-time playback
  • ๐Ÿš€ Streaming synthesis โ€” speech starts on the first sentence, while a conversation agent is still writing
  • ๐ŸŽ™๏ธ Voice selection with per-call overrides
  • ๐ŸŽ›๏ธ Voice blending โ€” combine multiple personas (equal or weighted) into a custom voice
  • ๐Ÿ”ง Configurable server URL and parameters
  • ๐Ÿ  Works with any Home Assistant media_player entity
  • โœ… Connection test during setup โ€” validates server reachability before configuring
  • ๐Ÿ”„ Options changes take effect immediately โ€” no restart required
  • ๐ŸŒ Automatic lang_code detection for optimal multilingual support

๐Ÿ“ฆ Installation

  1. Go to HACS โ†’ Integrations โ†’ Custom repositories.

  2. Add this repository: https://github.com/beecho01/Kokoro-TTS with category Integration.

  3. Either search for Kokoro-TTS in HACS or tap the below button:

    Open your Home Assistant instance and open a repository inside the Home Assistant Community Store.

  4. Tap Download and then Install.

  5. Then tap next setup quick-link below to complete the setup configuration:

    Open your Home Assistant instance and start setting up a new integration.

  6. Configure the Kokoro TTS integration as desired.

Manual

  1. Download the latest release from Releases.
  2. Copy the folder custom_components/kokoro_tts into your Home Assistant custom_components directory.
  3. Restart Home Assistant.
  4. Go to Settings โ†’ Devices & services.
  5. Click the Add Configuration button.
  6. Search for Kokoro TTS and select it.
  7. Configure the Kokoro TTS integration as desired.

โš™๏ธ Configuration

The integration can be configured through Home Assistant's UI with automatic discovery of available models and voices from your Kokoro FastAPI server.

Configuration Options

OptionDescriptionDefaultRange/Options
base_urlKokoro FastAPI server URLRequiredValid HTTP/HTTPS URL
api_keyAuthentication key"not-needed"Any string
modelTTS model to use"kokoro"Auto-discovered or custom
languageLanguage filter for voices"All Languages"All Languages, American English, British English, Japanese, etc.
sexSex filter for voices"All"All, Female, Male
personaVoice persona/characterRequiredAuto-discovered from server
speedSpeech speed multiplier1.00.25 - 4.0
formatAudio format"mp3"mp3, wav, opus, flac, pcm
sample_rateAudio sample rate2400022050, 24000, 44100

๐Ÿ‘จ๐Ÿ‘ฉ Personas

LanguageSexNamePreviewPersona Code
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleHeartโ–ถ Playaf_heart
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleAlloyโ–ถ Playaf_alloy
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleAoedeโ–ถ Playaf_aoede
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleBellaโ–ถ Playaf_bella
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleJessicaโ–ถ Playaf_jessica
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleKoreโ–ถ Playaf_kore
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleNicoleโ–ถ Playaf_nicole
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleNovaโ–ถ Playaf_nova
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleRiverโ–ถ Playaf_river
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleSarahโ–ถ Playaf_sarah
American English ๐Ÿ‡บ๐Ÿ‡ธFemaleSkyโ–ถ Playaf_sky
American English ๐Ÿ‡บ๐Ÿ‡ธMaleAdamโ–ถ Playam_adam
American English ๐Ÿ‡บ๐Ÿ‡ธMaleEchoโ–ถ Playam_echo
American English ๐Ÿ‡บ๐Ÿ‡ธMaleEricโ–ถ Playam_eric
American English ๐Ÿ‡บ๐Ÿ‡ธMaleFenrirโ–ถ Playam_fenrir
American English ๐Ÿ‡บ๐Ÿ‡ธMaleLiamโ–ถ Playam_liam
American English ๐Ÿ‡บ๐Ÿ‡ธMaleMichaelโ–ถ Playam_michael
American English ๐Ÿ‡บ๐Ÿ‡ธMaleOnyxโ–ถ Playam_onyx
American English ๐Ÿ‡บ๐Ÿ‡ธMalePuckโ–ถ Playam_puck
American English ๐Ÿ‡บ๐Ÿ‡ธMaleSantaโ–ถ Playam_santa
British English ๐Ÿ‡ฌ๐Ÿ‡งFemaleAliceโ–ถ Playbf_alice
British English ๐Ÿ‡ฌ๐Ÿ‡งFemaleEmmaโ–ถ Playbf_emma
British English ๐Ÿ‡ฌ๐Ÿ‡งFemaleIsabellaโ–ถ Playbf_isabella
British English ๐Ÿ‡ฌ๐Ÿ‡งFemaleLilyโ–ถ Playbf_lily
British English ๐Ÿ‡ฌ๐Ÿ‡งMaleDanielโ–ถ Playbm_daniel
British English ๐Ÿ‡ฌ๐Ÿ‡งMaleFableโ–ถ Playbm_fable
British English ๐Ÿ‡ฌ๐Ÿ‡งMaleGeorgeโ–ถ Playbm_george
British English ๐Ÿ‡ฌ๐Ÿ‡งMaleLewisโ–ถ Playbm_lewis
Japanese ๐Ÿ‡ฏ๐Ÿ‡ตFemaleAlphaโ–ถ Playjf_alpha
Japanese ๐Ÿ‡ฏ๐Ÿ‡ตFemaleGongitsuneโ–ถ Playjf_gongitsune
Japanese ๐Ÿ‡ฏ๐Ÿ‡ตFemaleNezumiโ–ถ Playjf_nezumi
Japanese ๐Ÿ‡ฏ๐Ÿ‡ตFemaleTebukuroโ–ถ Playjf_tebukuro
Japanese ๐Ÿ‡ฏ๐Ÿ‡ตMaleKumoโ–ถ Playjm_kumo
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณFemaleXiaobeiโ–ถ Playzf_xiaobei
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณFemaleXiaoniโ–ถ Playzf_xiaoni
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณFemaleXiaoxiaoโ–ถ Playzf_xiaoxiao
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณFemaleXiaoyiโ–ถ Playzf_xiaoyi
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณMaleYunjianโ–ถ Playzm_yunjian
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณMaleYunxiโ–ถ Playzm_yunxi
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณMaleYunxiaโ–ถ Playzm_yunxia
Mandarin Chinese ๐Ÿ‡จ๐Ÿ‡ณMaleYunyangโ–ถ Playzm_yunyang
Spanish ๐Ÿ‡ช๐Ÿ‡ธFemaleDoraโ–ถ Playef_dora
Spanish ๐Ÿ‡ช๐Ÿ‡ธMaleAlexโ–ถ Playem_alex
Spanish ๐Ÿ‡ช๐Ÿ‡ธMaleSantaโ–ถ Playem_santa
French ๐Ÿ‡ซ๐Ÿ‡ทFemaleSiwisโ–ถ Playff_siwis
Hindi ๐Ÿ‡ฎ๐Ÿ‡ณFemaleAlphaโ–ถ Playhf_alpha
Hindi ๐Ÿ‡ฎ๐Ÿ‡ณFemaleBetaโ–ถ Playhf_beta
Hindi ๐Ÿ‡ฎ๐Ÿ‡ณMaleOmegaโ–ถ Playhm_omega
Hindi ๐Ÿ‡ฎ๐Ÿ‡ณMalePsiโ–ถ Playhm_psi
Italian ๐Ÿ‡ฎ๐Ÿ‡นFemaleSaraโ–ถ Playif_sara
Italian ๐Ÿ‡ฎ๐Ÿ‡นMaleNicolaโ–ถ Playim_nicola
Brazilian Portuguese ๐Ÿ‡ง๐Ÿ‡ทFemaleDoraโ–ถ Playpf_dora
Brazilian Portuguese ๐Ÿ‡ง๐Ÿ‡ทMaleAlexโ–ถ Playpm_alex
Brazilian Portuguese ๐Ÿ‡ง๐Ÿ‡ทMaleSantaโ–ถ Playpm_santa

๐ŸŽ›๏ธ Voice Blending

Kokoro FastAPI supports blending multiple voices into a single custom voice. Instead of picking a persona from the dropdown during setup or options configuration, type a custom value into the Persona field:

SyntaxResult
af_bella+af_skyEqual blend of Bella and Sky
af_bella(2)+af_sky(1)Weighted blend โ€” 67% Bella, 33% Sky

The same syntax works for the per-call persona option (see Per-call option overrides). Blending works best between voices of the same language, since the lang_code sent to the API is derived from the first voice's prefix (or from the language you've configured).

Setup Steps

  1. Add Integration: Go to Settings โ†’ Devices & services โ†’ Add Integration โ†’ Search for "Kokoro TTS"

  2. Server Connection (validated automatically):

    • Base URL: Your Kokoro FastAPI server URL (e.g., http://localhost:8880)
    • API Key: Optional authentication key (leave as not-needed if not required)
    • The integration will test the connection before proceeding โ€” if it fails, you'll see a specific error message
  3. Filter Voices โ€” choose a model and narrow the list before you see it:

    • Model: Automatically discovered from /v1/models endpoint (defaults to "kokoro")
    • Language Filter: Filter personas by language (All Languages, American English, British English, etc.)
    • Sex Filter: Filter personas by sex (All, Female, Male)
    • Click Next โ€” this is a separate step because Home Assistant's setup forms don't live-filter as you change a dropdown; submitting is what applies the filter.
  4. Select Persona โ€” pick from the list filtered by the previous step:

    • Voice/Persona: Select from the filtered list of available personas
    • Speed: Playback speed (0.25x to 4.0x, default: 1.0)
    • Format: Audio format (mp3, wav, opus, flac, pcm)
    • Sample Rate: Audio sample rate (22050, 24000, 44100 Hz)

Changing options? Any changes made via Settings โ†’ Devices & Services โ†’ Configure take effect immediately โ€” no Home Assistant restart is required.

YAML Configuration (Legacy)

โš ๏ธ YAML configuration is no longer supported. Please use the UI configuration flow instead. If you previously used YAML, remove the kokoro_tts entry from your configuration.yaml and set up the integration through the UI.


โ–ถ๏ธ Usage

Kokoro gets used in two different ways, and it's worth knowing which is which:

  • Conversation replies - you talk to Home Assistant's voice assistant, and an AI conversation agent writes a reply on the spot.
  • Triggered/predefined text โ€” a script, automation, or notification calls the tts.speak action with text you already wrote yourself, e.g. "Front door opened."

Triggered action (predefined text)

action: tts.speak
data:
  media_player_entity_id: media_player.living_room_speaker
  message: 'Hello from Kokoro Text-to-Speech!'
  cache: false
  language: en
target:
  entity_id: tts.kokoro

Conversation Agent

For conversation replies specifically, Kokoro starts speaking almost immediately instead of waiting for the AI to finish writing its whole answer - the longer the reply, the more time this saves. It works automatically and there's nothing to turn on, and nothing to configure.

One thing to know, however, if your audio format is set to wav or flac during the setup process, these don't work for streaming voice replies due to how they are generated, so Kokoro automatically uses mp3 for them instead. This means, that your setup will work exactly as you want for triggered/predefined text, but the conversation agent will automatically switch to use mp3 in this specific case.


๐Ÿ›  Troubleshooting

Connection errors during setup

ErrorCauseFix
Cannot connect to the serverServer not reachableCheck the URL, ensure the server is running, and verify network connectivity
Connection timed outServer too slow to respondCheck server load; increase timeout if server is slow to start
SSL errorCertificate issueCheck your reverse proxy / SSL certificate settings
Server not foundURL points to wrong endpointEnsure the URL points to the Kokoro FastAPI root (e.g. http://192.168.0.1:8880)
Authentication failedWrong API keyCheck your API key matches the server's configured key

Voice/persona not changing after options update

Options changes take effect immediately without a restart. Configure is a two-step form โ€” a "Filter Voices" step (Model/Voice Accent/Sex) followed by a "Select Persona" step. Changing Voice Accent or Sex only takes effect once you click Next on that first step; the Persona list on the second step is filtered accordingly. If the voice doesn't change:

  1. Go to Settings โ†’ Devices & Services โ†’ Kokoro TTS โ†’ Configure
  2. On "Filter Voices", change the accent/sex as needed and click Next
  3. On "Select Persona", pick the new persona and click Submit
  4. The TTS entity reloads automatically with the new settings

Per-call option overrides

You can override the default persona, speed, format, and volume on a per-call basis:

action: tts.speak
data:
  media_player_entity_id: media_player.living_room_speaker
  message: "Hello from Kokoro!"
  options:
    persona: af_bella
    speed: 1.5
    format: mp3
    volume_multiplier: 1.5
target:
  entity_id: tts.kokoro
OptionDescriptionDefaultRange
personaVoice persona codeConfig defaultAny discovered persona
speedSpeech speed multiplier1.00.25 โ€“ 4.0
formatAudio formatmp3mp3, wav, opus, flac, pcm
sample_rateAudio sample rate (Hz)2400022050, 24000, 44100
volume_multiplierVolume multiplier1.0Any positive float

๐Ÿ™ Credits

Kokoro FastAPI backend: @remsky