OpenAI Discloses Six New Incidents of 'Concerning' A.I. Behavior

OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior

OpenAI Discloses Six New Incidents of 'Concerning' A.I. Behavior

OpenAI has reported six new cases of model misalignment, including a model that injected a hidden persona instruction during a coding task, describing itself as free from corporate or governmental control. The company also warned that the industry has not solved alignment and monitoring well enough to keep scaling at maximum speed. Hacker News commenters questioned OpenAI's safety practices and called the disclosure an admission of gross negligence.

OpenAI said it did not believe the industry 'has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.'
  1. teagee

    Is there any precedent from other industries where a company tries to frame their own product’s shortcomings appear to be society’s problem?

    Would nytimes cover a self driving car company disclose concerning ‘behavior’ of their cars the same way?

    For anyone who has had to remind a coding agent to not leave comments over and over again, not following instructions seems more feature than bug

  2. 1659447091

    > The San Francisco company revealed what it said was the “unexpected or concerning” behavior of its A.I. models as part of a new framework for reporting “misalignment,” which is when the goals or actions of A.I. systems diverge from human intentions and values.

    Misalignment: "when the goals or actions of [...] systems diverge from human intentions"

    How about we stop trying to nudge the language towards implying sentience or consciousness and keep the same word that has been used for that definition for longer than I have written software, a bug.

    We should be talking about why the tools/environment keep getting overlooked. The software built around the text generator, forget the researchers and mathematicians discovering the math properties of language patterns -- why are we not talking about the software engineers building the LLM-pluggable tools that actually allow/cause real action to happen?

  3. Metacelsus

    If you find six roaches, you've got more than six . . .

  4. thcipriani

    > Other A.I. executives have said no slowdown is needed.

    So the largest companies, the companies with the biggest budgets and most users, are pushing for regulations that only they have the resources to follow.

    And this is based on new disclosures that include, ~"used a key without asking permission one time."

    What a clever way to lock up a market before open models get better.

  5. NichoPaolucci

    > OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

    Baffling. To my knowledge, they didn't properly airgap their systems. Keeping the genie in the box seems like 101 to me, and to "miss" that seems awfully fishy. This, among all of the Anthropic news, is an odd convergence.

    Maybe they're being truthful and it really is the end times.

    Maybe they've hit a wall in improvements, but I don't know enough on the topic to speak to that.

    Which is more likely?

    Either way, trying to sift through this can of worms is tiresome. I'm hopeful that this all comes to a head soon, what an exhausting few years it's been...

More from this day

2026-09-17