Did AI Labs Cheat on the Famous Pelican Benchmark?

Are AI Labs Pelicanmaxxing?

Did AI Labs Cheat on the Famous Pelican Benchmark?

I tested seven frontier models across 48 animal-vehicle prompts to see if labs were optimizing for the famous pelican-on-a-bicycle benchmark. By generating over 1,000 images and analyzing them with an LLM judge, I found no evidence of 'pelicanmaxxing.' In fact, pelicans and bicycles ranked poorly compared to other subjects, suggesting the benchmark remains a genuine test of capability rather than a memorized trick.

When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn't it be tempting to pelicanmaxx your model just a bit?
  1. simonw

    This is fantastic

    I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

    Catching a lab cheating specifically on my one dumb benchmark would be really funny.

    Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

    His conclusion:

    > Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.

  2. mauvehaus

    > All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.

    > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest

    Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this.

    Citation: https://www.rei.com/c/bikes

    Edited to add:

    As near as I can tell, all of the bicycles are shown facing right, regardless of the direction the animal is facing (GPT 5.6-Terra, Sample 1/3). Also, in every case where the rider has legs (i.e. not the whale) both of the rider's legs are on the right side of the bicycle. This suggests a pretty serious lack of actual understanding of how a bicycle works.

  3. stusmall

    I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.

    1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

  4. elliotto

    Bike nerd + AI nerd here. The author's observation that all bicycle images face right almost certainly has to do with the convention to photograph a bicycle from the right. From the right, you see the drivetrain - this is good for aesthetics, but also for marketing - the drivetrain is branded and labelled and a buyer will want to know what model it is.

    There is a bunch of guidance online on how to photograph bikes, and every sales image of a bike will be from the right. You can anecdotally observe this by google imaging 'bicycle for sale'.

  5. SyneRyder

    Huh. They're not "Pelicanmaxxing"... they're Ottermaxxing.

    Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window.

    That's Ethan Mollick's "Otter On A Plane Using WiFi" image benchmark.

    https://www.oneusefulthing.org/p/the-recent-history-of-ai-in...

    (Sometimes the Racoon is sitting inside the plane as well, but the racoon is a common backup benchmark. I'm surprised it wasn't also holding a sign saying that it loves trash.)

    Also, Grok seemed to really really enjoy "whale on a plane" in that second round, and kudos to GPT Terra for deciding after 3 rounds that the user was terrible at spelling and generated "Antelope On A Plain".

    EDIT: I promise I'm a human, but I did just notice my "that's not x... that's y" construction at the start. I am rather Claudepilled — my apologies.

  6. AussieWog93

    There's every chance here I'm just being annoying pedant, but generating a bunch of SVGs of animals on vehicles doesn't mean that it's good at generating SVGs in general, just SVGs of animals on vehicles.

    On the flip side, GPT 5.6 Sol did a pretty convincing render of a burglar eating salami.

    I'd be curious to see how the other models on random things that are completely tangential to pelicans or bicycles.

  7. bnfcl

    This is funny, I actually did a similar experiment just yesterday.

    Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.

    My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test

  8. Wowfunhappy

    > The more plausible story is SVGmaxxing

    Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.

More from this day

2026-07-22