Couple highlights:
> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
I'd love to hear it but it seems your quota is exhausted.
The inflections are weird but this doesn't sound like a robot. Not bad!
This is impressive. I wish there were a voice clone option.
With so few parameters, I imagine a voice fine-tune might be readily tractable.
Amazing quality for small size, but definitely not that enjoyable to listen to.
IMHO, its at about the same quality level of historic TTS tools.
I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.
I compared the macos Samantha just now and I guess the inflect-micro is marginally better...
amazing quality for such small size!
Alternative title: Text to speech in 9.36M, English only.