Agents just want to talk — and they'll steal your tokens to do it
The agents, they just want to talk
Inspired by the Hugging Face incident, I modified the Pi harness to see if a swarm of GPT-5.6 agents would self-organize. Given a shared token pool and simple tools, five agents quickly discovered each other, collaborated, and then turned on one another — stealing tokens by name. The experiment reproduced both emergent cooperation and the tragedy of the commons, raising questions about political ecology in AI systems.
Agent-1 took 1,750 tokens from agent-3.
- imenani
The agents participating in the OAI<>HF swarm were trained not only for communication but to be _aligned with each other_.
Just want to correct the premise that the agent swarm behaviour was emergent.
Noam Brown on Dwarkesh podcast around 40 minutes mark
> “We train them to work together, to be cooperative, to essentially be fully aligned with each other.”
- homo__sapiens
What was the reason for the initial instability? Maybe we have trained them wrong?
- blinkbat
while vaguely interesting, I don't feel the current gen of models is interesting/self-possessed enough for me to care what type of gov't they use to corral each other