Cerebras and OpenAI unveil GPT-5.6 Sol Ultrafast, delivering 750 tokens per second
Accelerating GPT-5.6 Sol Ultrafast

Cerebras and OpenAI have introduced Ultrafast Mode, a new service tier in the OpenAI API powered by Cerebras' Wafer-Scale Engine. It runs GPT-5.6 Sol at up to 750 output tokens per second without quality loss, making it 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode. In tests on Humanity's Last Exam, Ultrafast completed all 2,500 questions in 11 hours versus 78 hours for Claude Fable 5, and on GDP-Val it achieved a 5.6x end-to-end speedup. The service is now in limited preview.
Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch.