Hetzner Tests LLM Inference with Qwen on Its Own Infrastructure
Hetzner is working on LLM Inference

I tried Hetzner's new experimental LLM inference service, which offers a free, OpenAI-compatible API running on their own hardware. While currently limited to a single Qwen model with no SLA, the performance was surprisingly fast. This move suggests Hetzner might leverage its efficient data centers to turn spare GPU capacity into a low-margin inference product, though the lack of massive multi-GPU clusters remains a question.
If Hetzner keeps serving one or two smaller models, I do not really see it becoming an important inference provider. That would be a cool experiment, but not much more.