Thanks! Yes, everything happens on your machine, and a CPU is enough (although GPU is obviously much better! Both are supported. My Framework 16 runs this at 3.2x realtime on CPU).
How finetuning works:
1. mp3 recordings of your voice are stored on your disk (if you choose so).
2. A big Model (Whisper large-v3-turbo by default, configurable) re-transcribes them when when your PC is idle overnight.
3. A LoRA fine-tunes your base model.
`voiceio learn schedule on` runs all of this automatically. For repeated mistakes there's also an instant fix: `voiceio vocab add <word>` steers the decoder right away, and `voiceio corrections add wrong right` handles anything left.
Beyond Whisper you can run Parakeet or a whisper.cpp server (both experimental). Feedback appreciated, have tested this with a small set of people so far.