Mixedbread's Toast 1 matches frontier search at a fraction of the cost

Introducing Toast 1

Mixedbread's Toast 1 matches frontier search at a fraction of the cost

Mixedbread introduces Toast 1, a specialized search agent that matches or outperforms frontier models like Claude Opus 5 and GPT-5.6 Sol on search quality while being up to 10× cheaper and 12× faster. It works with any search backend but excels with Mixedbread Search. In benchmarks, it achieves state-of-the-art results on OfficeQA Pro V2 and cuts token usage by 3.5× in legal tasks. Toast 1 is available now via the Mixedbread API at discounted launch pricing.

Toast 1 frees up the context window of frontier models to let them spend their tokens on reaching the right answer.
  1. trjordan

    I deeply love this idea of specialized LLMs for search. It's also extremely confusing to me how rough Google's entrance here is.

    When I, a human, need an answer to anything moderately complex, it's unlikely that I get it on the first (pre-AI) round of google searching. Simple stuff, sure, but more likely I'll need to go 2-5 rounds. Maybe click a few links. Double-check my assumptions.

    An LLM that can do that quickly seems like a slam dunk. I wonder what other problems benefit from that 10x-100x increase in context + 2-5 rounds with the LLM.

  2. satvikpendem

    Looks good as I use something similar with the SearXNG MCP, but a shame this isn't an open weight model. There are some wrappers around SearXNG which seem to reduce the token counts returned thus making it easier for the calling model to understand, but a full dedicated model for search is nice. How does it compare with Perplexity, Gemini with search, and Parallel AI? Those are the cloud providers of search based models that I've seen so far.

  3. blitzar

    I really wanted this to be a hardware startup - the Juicero of toast.

    Sadly its another software company.

  4. andai

    Article should probably explain what "Mixedbread Search" is.

  5. tolugenius

    I guess someone who has used a search agent (or a dedicated subagent) can speak when I'd reach for a tool like this vs either just 1) a smaller general model or 2) a non-llm approach to the problem? Like it's interesting I'm just curious how a search agent compares to say a model with dedicated rag pipelines is that much different?

More from this day

2026-08-14