Hetzner Tests LLM Inference with Qwen on Its Own Infrastructure

Hetzner is working on LLM Inference

Hetzner Tests LLM Inference with Qwen on Its Own Infrastructure

I tried Hetzner's new experimental LLM inference service, which offers a free, OpenAI-compatible API running on their own hardware. While currently limited to a single Qwen model with no SLA, the performance was surprisingly fast. This move suggests Hetzner might leverage its efficient data centers to turn spare GPU capacity into a low-margin inference product, though the lack of massive multi-GPU clusters remains a question.

If Hetzner keeps serving one or two smaller models, I do not really see it becoming an important inference provider. That would be a cool experiment, but not much more.

More from this day

2026-07-24