Etched Sohu vs. NVIDIA: Transformer ASIC vs. GPU (2026)

Etched Sohu vs. Nvidia: Transformer ASIC vs. GPU (2026) – Spheron Blog

Etched Sohu vs. NVIDIA: Transformer ASIC vs. GPU (2026)

Etched AI's Sohu is a transformer-only ASIC that hard-codes attention into silicon, claiming 500,000 tokens/sec on an 8-chip server for Llama 70B—about 62,500 tokens/sec per chip, versus ~700 tokens/sec for an H100 at batch 1. But Sohu sacrifices all programmability: it can't run vision, diffusion, MoE, or SSM models. With $800M raised, $1B in contracts, and first racks shipping summer 2026, the real question is whether the architectural bet pays off for your workload. This analysis compares Sohu against H100, B200, and Groq's LPU, offering a cost-per-token framework and a decision guide.

Sohu's throughput advantage over GPUs comes from architectural specialization of transformer attention patterns built on top of standard HBM3E, not from a SRAM-based design like Groq.

More from this day

2026-08-23