Record and transcribe voice input when user wants to speak instead of type, describe complex issues verbally, provide audio input, or dictate text...
This skill enables local voice transcription using whisper.cpp for privacy-preserving speech-to-text.
Use this skill when the user:
The transcription script now includes:
If the script detects missing installation, it will return JSON with "installation_needed": true. When you see this:
Offer to run installation:
"It looks like VoiceType isn't fully installed. Would you like me to run the installer? I can do this with: /voicetype-install"
If user agrees, run:
bash install.sh
Or use the /voicetype-install command which provides guided installation.
The script automatically handles:
.whisper/bin/ if not runningYou don't need to manually check the server - the script does it!
Run the transcription script:
source venv/bin/activate && python skills/voice/scripts/transcribe.py --duration 5
The script automatically:
Parse the output:
{"text": "transcribed speech", "duration": 5}{"error": "...", "installation_needed": true, "missing_components": [...], "help": [...]}{"error": "error message", "help": [...]}Handle installation_needed:
If JSON contains "installation_needed": true:
User: "Let me record a voice note about the bug I'm seeing"
Assistant:
{"text": "The submit button isn't working when I click it on the checkout page"}User: "Record my voice"
Assistant:
{"error": "VoiceType is not fully installed", "installation_needed": true, "missing_components": ["Python venv", "whisper.cpp binary"]}/voicetype-install or bash install.shThe transcription script accepts optional parameters:
--duration N - Record for N seconds (1-30, default 5)python skills/voice/scripts/transcribe.py --duration 10If transcription fails:
Check microphone access:
python -c "import sounddevice as sd; print(sd.query_devices())"
Verify whisper server:
systemctl --user status whisper-server
journalctl --user -u whisper-server -n 20
Test the script directly:
cd /path/to/voicetype
source venv/bin/activate
python skills/voice/scripts/transcribe.py
All voice processing happens locally: