My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."

A developer's personal benchmark pits AI models against a single prompt: generate an SVG of a frog with a Habsburg jaw. Each model gets three tries a month. The site tracks 14 models and 42 runs, all producing SVGs. The results showcase how different models interpret the prompt, with some adding editorializing comments about the frog's exaggerated jaw.
The annotations are mostly structural labels, but include some editorializing about the jaw feature: "massive protruding mandible" and describing the upper lip as "recessed, tucked behind the jaw" and lower teeth as "protruding" over the upper lip, which offer anatomical interpretation beyond plain labeling.