Benchmarking 15 E-Waste GPUs with Modern AI Workloads

Benchmarking 15 "E-Waste" GPUs with Modern Workloads

Benchmarking 15 E-Waste GPUs with Modern AI Workloads

I spent a year testing decommissioned NVIDIA Tesla GPUs like the K80, P100, and V100 to see if they can handle modern AI tasks. Despite being end-of-life and power-hungry, these affordable cards run LLMs and computer vision models surprisingly well when paired with custom cooling and older software stacks. My goal was to build an inexpensive homelab node, proving that you don't always need the latest hardware to experiment with cutting-edge technology.

It is irresponsible to suggest these cards be used today... is the kind of finger wagging that has no place in homelabbing.
  1. SillyUsername

    No mention of the venerable Tesla P4. 75W peak, 8GB VRAM, about $80 (£60).

    I have 6x P4s, a Xeon E5 2696v3 (36 threads, 3.8ghz peak but all core turbo unlocked, so 6 cores at 3.8Ghz - about 8 cores at 3.5ghz, or all cores at 3.1ghz), 48GB DDR4, all fit into a micro atx case running on a 650W MSI psu.

    This gives me a virtual 48GB GPU (llama.cpp ftw) to backup that 48GB of RAM.

    I typically see scores of at least 7-12t/s on 20-30B Q4KM size dense models, on a 32K/48K/64K context, adequate for modern inference.

    The pain point is the prompt loading, it is far far slower, minutes not seconds, than modern tensor core 8GB 5060s (my other machine's 2x GPUs) but is quite similar in regular inference speed once it has loaded.

  2. SwellJoe

    When I wanted to tinker with self-hosted models, I bought a couple of Radeon Pro V620 GPUs, because they're 32GB, still supported by current ROCm releases, and a few years newer than the similar-priced 32GB Nvidia cards (which are all EOL). They're a little faster than the old Tesla stuff, as well. 64GB is enough to run Gemma 4 31b 4-bit QAT with pretty big context at a respectable interactive speed (30+ tokens per second sustained).

    That said, even the old Radeon Pro stuff has gotten more expensive on eBay, so I'm not necessarily recommending cheap old server cards that need custom-printed fan shrouds to operate in a consumer PC. Probably better to buy the Radeon AI Pro R9700 for $1400, which will be faster, supported for many years, and has a fan already. Or, maybe even the Intel ARC B70 for $1000.

  3. latchkey

    Darn, I was hoping to see bc-250's (aka PS5 chips) in there. They've recently become popular for inference and they are only about $200 on ebay. They hold a special place in my heart because I deployed 20k of them and I'm glad to see they are finding a purpose now and not just e-waste.

  4. kn100

    Great read. I'd love to know more about how power consumption changes as cards get newer too!

  5. roger_

    Getting (further) into this myself so good timing. Running Qwen 3.6 27B at decent speed on some old cards but going to branch out.

    I bough an Octominer for ~$150 which has power and PCIe slots and a basic Celeron and should let me expand to as many GPUs as I want.

    I considered the P100s but I think the V100 16GBs are a better deal at $250. The 32GBs are way too much though.

More from this day

2026-07-13