Inflect-Micro-v2: Complete Local Text-to-Speech Under 10 Million Parameters
Inflect-Micro-v2: complete voice in 9.36M parameters

I built and funded Inflect-Micro-v2 independently to deliver high-quality, fixed-voice English text-to-speech synthesis with fewer than 10 million parameters. This model runs efficiently on CPU or CUDA, handles long text, and offers deterministic seeds. With a 66.2% human preference rate and fast inference speeds, it proves that compact models can achieve professional audio quality without massive compute requirements.
No single metric captures TTS quality, so I report human preference, predicted naturalness, multi-ASR intelligibility, complete footprint, and runtime separately rather than compressing them into one unverifiable score.
- yjftsjthsd-h
Couple highlights:
> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
- modinfo
This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!
here my implementation with speech dispatcher and server:
https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!
- NetOpWibby
The inflections are weird but this doesn't sound like a robot. Not bad!