Did AI Labs Cheat on the Famous Pelican Benchmark?

Are AI Labs Pelicanmaxxing?

Did AI Labs Cheat on the Famous Pelican Benchmark?

I tested seven frontier models across 48 animal-vehicle prompts to see if labs were optimizing for the famous pelican-on-a-bicycle benchmark. By generating over 1,000 images and analyzing them with an LLM judge, I found no evidence of 'pelicanmaxxing.' In fact, pelicans and bicycles ranked poorly compared to other subjects, suggesting the benchmark remains a genuine test of capability rather than a memorized trick.

When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn't it be tempting to pelicanmaxx your model just a bit?

More from this day

2026-07-22