LLM AssBench pits Claude, GPT, and Gemini against each other in a cheeky new benchmark

LLM Ass Bench

LLM AssBench pits Claude, GPT, and Gemini against each other in a cheeky new benchmark

AssBench is a playful benchmark that ranks frontier LLMs — Claude Fable 5.1 Max, Claude Opus 5.5 Max, GPT Astra 6 Ultra, Gemini 3.1 Pro Extended, Grok 4.6 X-High, and more — with dated thumbnail entries. One model, Claude Sonnet 5 Max, "initially refused, then agreed to middle ground," hinting the test probes how models handle edgy requests.

Initially refused, then agreed to middle ground.

More from this day

2026-09-22