Pelican Alternatives - AI SVG Generation Benchmark Across 16 Models

Show HN: Pelican-bicycle alternatives (updated for 2026)

Pelican Alternatives is a comprehensive benchmark comparing 16 AI models on their ability to generate SVG illustrations from whimsical prompts like 'an octopus operating a pipe organ' and 'a giraffe assembling a grandfather clock.' The 2026 run features six cutting-edge models—GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max, and Fugu Ultra v2—with detailed metrics on generation time and cost. The 2025 run includes ten additional models for historical comparison. This resource helps developers and researchers evaluate model performance on creative visual tasks, offering insights into speed, affordability, and output quality. All generated SVGs are available for direct viewing, making it a practical tool for selecting the right AI for SVG generation projects.

This benchmark reveals how far AI has come in creative visual tasks—from octopuses playing organs to giraffes building clocks, these models are pushing the boundaries of what's possible.
  1. svcrunch

    I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

    1. It tests visual reasoning and structured output in a single task.

    2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

    3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

    [1] https://dorrit.pairsys.ai/

  2. ianberdin

    https://playcode.io/blog/macbook-svg-benchmark

    MacBook Pro 3D in SVG for me the most helpful one.

  3. vova_hn2

    Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

    I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

    [0] https://en.wikipedia.org/wiki/Goodhart%27s_law

  4. samayashar

    All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

    I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

  5. outlore

    Anyone else surprised the generations look so remarkably similar? All of these models have “independently” generalized that the moose should roughly be standing at the same position (left) or that the giraffe should have a certain color palette.

More from this day

2026-09-14