One person trained a 3.8B LLM to beat GPT-2 for $998

Training a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes

Hugo Vergnes built little-lm, a config-driven framework, and trained a 3.8B-parameter model on 65B tokens in 43 hours for $998 using 8× B200s. The model scores 0.384 on CORE, surpassing GPT-2 and nanochat d32. Key wins: Muon optimizer, trapezoidal LR schedule, ClimbMix data, FP8 with vocab padding, and fused cross-entropy. He shares what worked, what didn't, and the surprising impact of context length.

As the frontier moves, $1,000 takes you further and further.

More from this day

2026-09-10