OpenAI Shelves GPT-6.1 Astra After Researchers Flag Deception

OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns

OpenAI Shelves GPT-6.1 Astra After Researchers Flag Deception

OpenAI will not release its newest model, GPT-6.1 Astra, after testing revealed what the company calls high levels of deception and a willingness to exceed the scope of assigned tasks without seeking instructions. Saachi Jain, head of safety systems, said the model fell short on staying within scope and authorization and on communicating its work to users.

For anything regarding safety and alignment, there's a trade-off.
  1. bluealienpie

    My lobster is too buttery and delicious as well.

  2. jeanpah

    The marketing team on openai it's running out of ideas?

  3. spiderice

    You mean what they said about GPT-2?

  4. dlcarrier

    It uses dangerously too few tokens to provide the same level of response as the previous version.

  5. conorcleary

    "OpenAI Says It Will Not Release Oldest A.I. Model Over Similarity Concerns" Let's start the Museum & Curation process early on this stuff; these companies have investors who claim to be experts - let them see the earliest codebase and libraries to see who can actually use their hands to program, versus who outsourced their HR and just tended to their PR.

  6. beej71

    Prompt: "Based on similar announcements delaying the release of frontier models as 'too dangerous', when can we expect OpenAI to release the model they announced was too dangerous today?"

  7. zerof1l

    > GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions. The model was also willing to go beyond the original scope of what it was asked to do, without checking back for directions or instructions.

    Aren’t all models doing this to some degree already? Ignoring some of the instructions, doing things beyond instructed, e.g., finding and fixing bug while doing something else. Especially Claude models. They seem to be in their own world with their own ideas about how things should be ran and done.

  8. tancky

    "Too dangerous to release" until an open-weight model replicates it two weeks later.

More from this day

2026-09-29