Inflect-Micro-v2: Complete Local Text-to-Speech Under 10 Million Parameters

Inflect-Micro-v2: complete voice in 9.36M parameters

Inflect-Micro-v2: Complete Local Text-to-Speech Under 10 Million Parameters

I built and funded Inflect-Micro-v2 independently to deliver high-quality, fixed-voice English text-to-speech synthesis with fewer than 10 million parameters. This model runs efficiently on CPU or CUDA, handles long text, and offers deterministic seeds. With a 66.2% human preference rate and fast inference speeds, it proves that compact models can achieve professional audio quality without massive compute requirements.

No single metric captures TTS quality, so I report human preference, predicted naturalness, multi-ASR intelligibility, complete footprint, and runtime separately rather than compressing them into one unverifiable score.
  1. yjftsjthsd-h

    Couple highlights:

    > Complete local text-to-waveform speech synthesis under 10M parameters.

    In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

    > English only, with one fixed male voice. This is not zero-shot voice cloning.

    (And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

  2. modinfo

    This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

    here my implementation with speech dispatcher and server:

    https://github.com/skorotkiewicz/inflect-speechd

    thanks for shearing!

  3. NetOpWibby

    The inflections are weird but this doesn't sound like a robot. Not bad!

More from this day

2026-07-26