CrucibleBench: Using a MUD to Evaluate LLM Behavior for $99

Can a MUD evaluate LLMs? A $99 proof of concept

CrucibleBench: Using a MUD to Evaluate LLM Behavior for $99

I built CrucibleBench, a proof-of-concept using a persistent MUD to test how language models handle trust and social objectives over 50 turns. By placing models in a constrained text world where mistakes leave traces, we identified specific failure modes like dialogue looping. Our $99 experiment revealed that LLM judges can drastically reorder rankings, proving that aggregate reliability statistics often hide critical measurement instability.

We did not choose a MUD because it is charming. We chose it because its constraints make behavior measurable.

More from this day

2026-07-22