Did AI Labs Cheat on the Famous Pelican Benchmark?
Are AI Labs Pelicanmaxxing?

I tested seven frontier models across 48 animal-vehicle prompts to see if labs were optimizing for the famous pelican-on-a-bicycle benchmark. By generating over 1,000 images and analyzing them with an LLM judge, I found no evidence of 'pelicanmaxxing.' In fact, pelicans and bicycles ranked poorly compared to other subjects, suggesting the benchmark remains a genuine test of capability rather than a memorized trick.
When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn't it be tempting to pelicanmaxx your model just a bit?