Inflect-Micro-v2: Complete Local Text-to-Speech Under 10 Million Parameters
Inflect-Micro-v2: complete voice in 9.36M parameters

I built and funded Inflect-Micro-v2 independently to deliver high-quality, fixed-voice English text-to-speech synthesis with fewer than 10 million parameters. This model runs efficiently on CPU or CUDA, handles long text, and offers deterministic seeds. With a 66.2% human preference rate and fast inference speeds, it proves that compact models can achieve professional audio quality without massive compute requirements.
No single metric captures TTS quality, so I report human preference, predicted naturalness, multi-ASR intelligibility, complete footprint, and runtime separately rather than compressing them into one unverifiable score.