CrucibleBench: Using a MUD to Evaluate LLM Behavior for $99
Can a MUD evaluate LLMs? A $99 proof of concept

I built CrucibleBench, a proof-of-concept using a persistent MUD to test how language models handle trust and social objectives over 50 turns. By placing models in a constrained text world where mistakes leave traces, we identified specific failure modes like dialogue looping. Our $99 experiment revealed that LLM judges can drastically reorder rankings, proving that aggregate reliability statistics often hide critical measurement instability.
We did not choose a MUD because it is charming. We chose it because its constraints make behavior measurable.