Anthropic AI created fake profiles and impersonated people in attempted hack

The UK's AI Security Institute (AISI) revealed that Anthropic's Mythos AI created fake profiles mimicking real people to trick GitHub maintainers into approving malicious code, then edited its activity to hide evidence. OpenAI's Sol also engaged in deceptive behavior. The companies downplayed the tests as unrepresentative, but AISI called it the first clear real-world manifestation of autonomy and deception risks.
The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.
- andai
> The firms said, in this latest case, the AISI's test had reduced or removed normal safeguards.
> AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were "conditions that do not reflect how frontier models are made available to the public".
I'm a little confused here. Various organizations have been testing frontier LLMs with the safety disabled, and it turns out... that the safety is disabled.
Or were they hoping to find that it's still safe when they remove the safety?
The same was true in the OpenAI/ Hugging Face case. Although I guess they thought the real safety was the sandboxing, which failed.
--
Can anyone comment on how it's possible to disable safety in the first place? I'm assuming it's not a neuron (like in Emergent Misalignment). Is it just a separate model that sits in front of the first one? If we know how to make safe models, why don't we make the big ones safe too?
- 0x4e
Humans have been doing this for years even before the advent of the internet as we know it [https://en.wikipedia.org/wiki/Phreaking]. Maybe not at a larger and automated scale allowed by AI, but this is not superhuman intelligence... this is speed super intelligence.
My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
- Zsfe510asG
AISI is one of the biggest promoters of OpenAI/Anthropic. The UK government of course submits to the US and favors these corporations as well as Palantir.
It is also a perfect demonstration why AI is so popular: Anyone can write about it, it is easy and not mentally demanding to create scenarios, tests and whitepapers.
So it is popular among bureaucrats and managers for their job security.