Voice Dictation for Linux that learns as you speak

5 pointsposted 8 hours ago
by montenegrohugo

4 Comments

montenegrohugo

8 hours ago

I use voice dictation for 90% of my work and there was nothing good on Linux. So I built my own.

It gets better as you use it, finetuning on your recordings overnight. Fully local, no data leaves your PC, source available, free.

I believe the future of software is software you actually own and that adapts to your personal needs.

github: https://github.com/Hugo0/voiceio

ankitg12

8 hours ago

Nice - was looking for something like this after Whisper medium.en on Voxtype has repeatedly been making same mistakes for me. Would try to try it out :). In the meantime, do you have any writeup on how it does fine tuning? Does it happen locally on machine - needs CPU or GPU etc?

montenegrohugo

8 hours ago

Thanks! Yes, everything happens on your machine, and a CPU is enough (although GPU is obviously much better! Both are supported. My Framework 16 runs this at 3.2x realtime on CPU).

How finetuning works:

1. mp3 recordings of your voice are stored on your disk (if you choose so).

2. A big Model (Whisper large-v3-turbo by default, configurable) re-transcribes them when when your PC is idle overnight.

3. A LoRA fine-tunes your base model.

`voiceio learn schedule on` runs all of this automatically. For repeated mistakes there's also an instant fix: `voiceio vocab add <word>` steers the decoder right away, and `voiceio corrections add wrong right` handles anything left.

Beyond Whisper you can run Parakeet or a whisper.cpp server (both experimental). Feedback appreciated, have tested this with a small set of people so far.

user

8 hours ago

[deleted]