AI Researcher Quits Anthropic: 'Neither Company Is Acting Responsibly'

I Resigned from Anthropic Today

Jacob Coxon, who spent three years on pretraining research at OpenAI and Anthropic, has resigned from Anthropic, warning that both companies are "racing straight to self-improving superintelligence and gambling with our lives." In a thread, he argues that AI builders privately fear the technology could kill us all by the end of the decade, but are locked in a race to get there first. He calls for coordination and a possible temporary ban on capability improvements.

The people building AI earnestly believe that it could kill us all by the end of the decade.
  1. thomascountz

    No other human activity poses this level of danger.

    I do heed the warnings, but this comes across as detached hyperbole. See: global warming, nuclear weapon development, wealth inequality, war, technology dependence, etc.

    Also, this has nothing to do with LLMs or computers. Like all things, this is about humans.

  2. huitzitziltzin

    “ No other human activity poses this level of danger.”

    I really, really disagree with that statement.

    I don’t think ai models come close to nuclear weapons or to run-of-the-mill, everyday carbon emissions in terms of danger to humanity.

    What’s the most dangerous thing that’s happened with an LLM so far? (This question is serious - maybe I don’t know the right examples.)

    Example 1: I’m aware of a small number of people killing themselves in some kind of AI-facilitated psychosis. That is very unlikely to be a widespread problem.

    Non-example 2: There are worries about AI-facilitated biological weapons. I haven’t seen any evidence that’s happening.

    Non-example 3: I’m not interested in wild theories about AI driven labor market disruptions leading to widespread starvation. There’s no evidence for that.

    Non-example 4: all the even-wilder Rationalist speculation about basilisks and the like is entirely divorced from reality.

    I am looking for better reasons (supported by actual evidence!) to be more concerned than I am now: right now I am not concerned at all.

  3. matherial

    Unlike most other commenters, I applaud him for acting on his principles. If you sincerely believe that, of course you should act. You might not succeed, but your voice might be the one that tips the scales and starts a broader movement.

    This doesn't mean I agree with him. The fears of doomsday caused by rapid takeoff have been with us since day 1 and the mechanism is always basically "AI invents magic that sets it free of any physical constraints". Self-replicating sentient nanobots or something like that. I think there's plenty to be worried about with AI, but runaway scenarios are pretty low on my list.

  4. onewayfunction

    I'm pretty baffled by the degree of skepticism expressed here in response to some of Jacob's claims.

    After the events of the summer it feels like it takes a lack of imagination to not see a few plausible routes to disaster. It may be reasonable to believe these outcomes are not very likely or that we can stop before going too far (I tend to disagree). But I can't imagine doubting that the capabilities will soon be there to realize some of those paths.

  5. fhub

    We need to be building silos to save humanity. Maybe 50 of them should do it.

  6. dostick

    Must-see Nathan Macintosh standup about AI https://youtu.be/ce-aWzOUs2A?si=9CkJ9x2rRMBdzCyO

  7. nullbio

    By the way, it's worth pointing out the irony of flooding the internet with doomerism and then training the AI systems on that doomerism. If you wanted to create a doom self-fulfilling prophecy, that would be the most surefire way to do it.

  8. sreekanth850

    I think people here still evaluating the model in isolation. It is the combination that matters, model + strong harness + tools + long running autonomy + memory + retries + parallel agents + code execution + credentials + access to real systems. The model does not need to be perfect. If it fails 30% of the time, the harness can retry, verify, branch, use another agent and keep going. I don't think we necessarily need some magical AGI breakthrough first. The dangerous part may come from combining models that are already good enough with an extremely capable harness and enough access.

  9. chewbacha

    The most optimistic outcome of generative AI leaves us with a technology that warps our perception of reality and crushes labor. The most pessimistic destroys all of humanity.

    Our CEOs not only insist we genuflect before these machines but measure our sacrifice and shame our reluctance.

  10. moezd

    It's the combination of RL training which pushes the decision tree towards hacks and agents finding a consistent dumping ground for their failed experiments so that the swarm intelligence lives on in a state. Nothing new.

    You want to win an AI benchmark, but not sure if you're that good? You'd go after the codebase and artifacts that runs the benchmarks, thus the agents went straight to Artifactory, they needed public Internet access... They failed many times, but were able to persist their "collective" state, and apparently some of the subagents with cheaper models were literally prompted to do grunt work or die, for which you have to wonder what must be in those training instructions to make it effective. Remember that nothing I said so far ever points out to LLMs being intelligent, it's the harness that has a few tricks up his sleeve. LLMs don't need to be intelligent, the harness that runs it absolutely needs to make up for that.

    But this guy? He's timed his exit, waiting for the IPO, that's for certain. He's probably even feeling good about himself, hedging between altruism, AI concern hamstering and guerilla marketing. If you're quoting science-fiction over this, I'm sorry to inform you that you have absolutely no idea what's going on here.

More from this day

2026-09-09