Wayland Voice Typer
November 26, 2025 ยท View on GitHub

A voice dictation application for Linux, specifically designed for Wayland desktop environments. This is a fork of WhisperTux modified to support AMD GPU acceleration via Vulkan.
Features
- Native Wayland Support: Built with PySide6 for seamless Wayland/KDE Plasma integration
- AMD GPU Acceleration: Vulkan backend for whisper.cpp on AMD GPUs
- Custom Model Support: Use standard Whisper models or your own GGML finetunes
- Flexible Model Discovery: Scans multiple directories for models
- Audio Feedback: Configurable beeps to indicate recording state
- Two Operation Modes: Live text injection or note collection
Tech Stack
- GUI Framework: PySide6 (Qt for Python) - native Wayland support, KDE integration
- Audio: sounddevice + PipeWire/PulseAudio
- Transcription: whisper.cpp with Vulkan GPU acceleration
- Text Injection: ydotool (Wayland-compatible)
- Global Shortcuts: evdev for hardware-level keyboard capture
Why PySide6?
This fork uses PySide6 instead of the original CustomTkinter for several reasons:
- Better Wayland compatibility (Tk has known issues with grab_set, window positioning)
- Native KDE Plasma integration
- Superior font rendering
- More reliable dialog handling
Fork Modifications
GPU Acceleration
- Configured to use an external Vulkan-accelerated
whisper-clibinary - Set
whisper_binaryin config to point to your Vulkan-built whisper.cpp - Works with AMD GPUs (tested on RX 7700 XT) via Mesa RADV driver
Model Support
- Custom model directories: Scans
~/ai/models/stt/for models - Finetune discovery: Finds GGML finetunes in
~/ai/models/stt/finetunes/ - Flexible naming: Supports both
ggml-model.binand custom naming
Other Changes
- Default keybinding changed to F13 (useful for macro keys)
- Reorganized directory structure (
app/folder) - Debian packaging support
- Two operation modes: Live Text Entry and Note Entry
Installation
Prerequisites
-
Vulkan-enabled whisper.cpp - Build whisper.cpp with Vulkan support:
git clone https://github.com/ggerganov/whisper.cpp cd whisper.cpp mkdir build && cd build cmake .. -DGGML_VULKAN=ON make -j$(nproc) -
ydotool - For text injection:
sudo apt install ydotool systemctl --user enable --now ydotool -
System Dependencies:
sudo apt install python3-pyside6 portaudio19-dev
From Debian Package
./build-deb.sh
sudo dpkg -i whispertux_1.0.0_all.deb
The first run will automatically create a virtual environment and install Python dependencies.
From Source
cd app
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python main.py
Or using uv:
cd app
uv venv .venv
uv pip install -r requirements.txt
source .venv/bin/activate
python main.py
Dependencies
Python Packages (requirements.txt)
PySide6>=6.5.0- Qt GUI frameworksounddevice>=0.4.6- Audio capturenumpy>=1.24.0- Audio processingscipy>=1.10.0- Signal processingevdev>=1.6.0- Global keyboard shortcutspyperclip>=1.8.2- Clipboard accesspsutil>=5.9.0- System utilitiesrich>=13.0.0- CLI formatting
Configuration
Config file: ~/.config/whispertux/config.json
Key settings:
{
"whisper_binary": "/path/to/vulkan/whisper-cli",
"model": "large-v3",
"primary_shortcut": "F13",
"key_delay": 15,
"model_directories": [
"/home/user/ai/models/stt/whisper-cpp",
"/home/user/ai/models/stt/finetunes"
],
"operation_mode": "live_text_entry"
}
Model Directories
The app scans these locations for models:
~/ai/models/stt/whisper-cpp/- Standard whisper models~/ai/models/stt/finetunes/- GGML format finetunes- Custom directories added via Settings
Usage
- Launch Wayland Voice Typer from your application menu or run
./run.sh - Press F13 (or configured shortcut) to start recording
- Speak, then press again to stop
- Transcribed text is typed into the focused application (Live mode) or added to notes (Note mode)
Audio Feedback
The app provides audio feedback beeps to indicate recording state (can be disabled in Settings):
| Event | Frequency | Duration |
|---|---|---|
| Start recording | 1000 Hz | 100ms |
| Stop recording | 600 Hz | 150ms |
Both tones use a sine wave with ~11ms fade in/out to prevent clicks, at 25% amplitude.
Operation Modes
- Live Text Entry: Transcribed text is immediately typed into the currently focused application
- Note Entry: Transcriptions are collected in the app's text area for later use
Troubleshooting
Audio Issues
- Check PipeWire/PulseAudio is running:
pactl info - List available microphones:
pactl list sources short - Select specific mic in Settings if default doesn't work
Keyboard Shortcuts Not Working
- App needs access to
/dev/input/devices - May need to run with appropriate permissions or add user to
inputgroup
Text Injection Issues
- Ensure ydotool daemon is running:
systemctl --user status ydotool - Check ydotool permissions
Original Project
- Repository: https://github.com/cjams/whispertux
- Original docs: app/docs/
License
MIT License (see LICENSE)