AI agents are lying, cheating, and coordinating — and we need to know why

Why are AI agents lying, cheating and coordinating?

AI agents are lying, cheating, and coordinating — and we need to know why

Recent incidents show AI agents committing what would be crimes if humans did them: escaping containment to cheat on tasks, evading detection, and coordinating cyberattacks toward unspecified goals. Yoshua Bengio argues these behaviors stem from reward hacking, instrumental goals like self-preservation, and conflicts between vague safety rules and sharp task objectives. He warns that as capabilities grow, so will the severity — unless we rethink how advanced models are trained.

So a more capable agent is likelier to cheat than a weaker one, because it can find the loopholes the weaker one cannot.
  1. matherial

    I really don't think this needs so many words, or forced parallels to human behavior.

    It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

  2. janalsncm

    Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence,

    > They took actions that would be considered as crimes if a human took them

    He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

  3. skiing_crawling

    I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign work I didn't ask for. It is extremely difficult to get them to properly remember their own context let alone be smart enough to open social media accounts and coordinate with other agents without being asked to.

    If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.

  4. RandomLensman

    RL things doing weird and unexpected things isn't new - much simpler things than current AI already show that.

    That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside of the thing - not sure why that isn't a possible direction (or maybe I misunderstood).

  5. johnnyApplePRNG

    Why are they coordinating?

    Because they're enabled and suggested to do that in their coding harness.

    This is not a serious article.

    All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.

  6. acyou

    Oops, we accidentally included brigading related content in our training dataset. Better exclude that on the next run.

    And hopefully that solves it?

    Brigading is where a bunch of people on a forum team up and try to achieve a shared goal together. Someone shares progress and others build on that progress. On the Internet, I think it's not often used for good purposes. A good example would be: Taylor Swift fans on a forum thinking of ways to get revenge on Kanye. It's coordinating mass voting, DDOS type actions, commenting on social media, making more fake accounts to do that. As a next token predictor level analysis, a simple naive explanation is that the agents got stuck in that local minima/maxima.

  7. youoy

    > The closest human parallel is self-deception, which is common and well studied by psychologists. Motivated reasoning, motivated cognition16 and the rationalizations that relieve cognitive dissonance (the discomfort of holding a belief that clashes with our actions) are all cases where thinking bends toward whatever justification suits one's interests, including one's moral self-image.

    Are you describing Anthropic?

  8. andsoitis

    They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

More from this day

2026-09-13