LoRA Speedrun: The Public Leaderboard for Fine-Tuning Techniques

LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques

LoRA Speedrun: The Public Leaderboard for Fine-Tuning Techniques

I created a public wall-clock leaderboard to race fine-tuning techniques on a frozen task and hardware. Using Qwen2.5-1.5B and GSM8K on a single NVIDIA L40S, anyone can compete for free via Modal. We verify every record with three fresh runs to ensure fair, apples-to-apples comparisons, turning parameter-efficient fine-tuning into a transparent, scientific arena.

LoRA/QLoRA is how most real-world fine-tuning actually happens, and the technique space exploded, but there's no adversarial, apples-to-apples arena where these ideas race each other in public.
  1. jmward01

    I think there is real value in going smaller/limiting resources. The trend is 'just make the weights bigger and throw more data at it'. It is a MBA's view of winning. We have a knob, keep turning it. It does work but it may not drive as much creativity as resource limits can drive. It is like urban growth boundaries in city planning. If you aren't allowed to 'just expand' you are forced to build more intelligently inside the city and those creative solutions often lead to major improvements.

  2. patrick0d

    Inspired by parameter golf and speedrun approaches I make the case for picking loss functions like a wallclock for LoRA on AI safety targets. The result when I tried it was a functional distillation of an Sparse AutoEncoder into a 5.3MB probe. I have a technical writeup below about it if anyone is interested.

    https://www.lesswrong.com/posts/PagGF8roBJmjLunsX/competitiv...

  3. stephantul

    Sadly 100% generated.

    I think the idea is interesting though, although I wonder if training time for LoRA is such a bottleneck to deserve its own, extremely narrowly scoped, leaderboard. Maybe if it was more tasks or more models we could hope that it transfers? With a single task, and a single model, I’d be afraid of this overfitting pretty heavily.

    For NanoGPT, I think the idea always was that the ideas can be transferred to much larger models, or serve as stepping stones for investigations on larger models.

More from this day

2026-07-20