OpenAI's GPT-6 Astra page returns 404, leaving users with a poetic error message

OpenAI's GPT-6 Astra announcement page is returning a 404 error, but the error page itself features a cryptic poem: 'Signal drifts past Mars / Ground lights wait with steady hands / Stars answer in time,' attributed to 'gpt-5.6-sol.' The page suggests GPT-6 Astra may be an upcoming model, but OpenAI has not yet made an official announcement.

Signal drifts past Mars Ground lights wait with steady hands Stars answer in time
  1. dang

    Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273

    How about we stick to that one for talking about the rollout, and this one for talking about the model?

  2. intenex

    The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.

    Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.

    I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.

    For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this […]

  3. abixb

    I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

    If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

    As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

  4. manlymuppet

    I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?

    Even if I did trust an AI to get everything right, it's not like the AI can read my mind.

    If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?

    All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.

    (Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)

  5. astrobiased

    I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547

    Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.

    It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.

    The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?

    With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.

  6. dalemhurley

    OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.

    Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).

    Codex is slightly better than Claude Code.

    Good on Sam Altman getting back to basics and turning OpenAI around.

  7. tristanj

    GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

    Performance is significantly higher than Fable 5.1

    Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

  8. datadrivenangel

    Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.

    Same for HealthBench Professional and a few others.

    Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.

  9. jumploops

    I think the thing I'm most excited about is the increase in _user prompting_.

    If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.

    The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.

    It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.

    Hopefully this model has the right balance, or at least better?

  10. XCSme

    It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?

More from this day

2026-09-03