Why I'm Still Bearish on LLMs After Navier-Stokes

Despite headline successes like solving Navier-Stokes problems in Lean, I argue that current LLMs remain far from autonomous knowledge workers. They need constant oversight, fail on slight task variations, and reward hacking persists. Rigorous specification is costly and rare, human review doesn't scale, and only three narrow firm types can adopt fully autonomous LLMs. Most will rely on cheap open models, not frontier labs.

For most domains LLMs will continue to look like a cracked intern: quick and effective in the hands of an adult but not given run of the place.
  1. carodgers

    This April 2026 paper is a fun and related read.

    https://arxiv.org/html/2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.

  2. keeda

    The premise in the very first point seems off:

    > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers...

    Even assuming this is how the AI companies are being valued (they're not), the numbers are off.

    The "value" of most knowledge workers -- based on what enterprises currently pay for them -- is $50 - 70 trillion annually. It's reasonable to assume that if AI drop-in-replaced all those knowledge workers, AI companies could credibly charge somewhere in that order of magnitude, because that's what the market is already bearing.

    So if their hypothetical revenues are double-digit trillions and valuations are some multiple of that, the entire AI industry would be valued at double-digit trillions at the least.

    Yet cumulatively the industry (the frontier labs + the SWAG estimate of the AI parts of all the other players) are valued at, say, ~6 - 7 trillion? Which seems like a fair approximation of how much knowledge work they can currently automate.

  3. knuppar

    Short and to the point! Open and cheap models will undercut the big labs continuously. The blast radius won't be pretty once spending commitments knock the door.

  4. bluegatty

    "are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, "

    No, they're really not.

    They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be.

    And that they will capture most of that ... which they won't.

    The Frontier Labs are a very bad buy at a high price, but that partly has to do with wacky pricing, but actually mostly has to do with their relatively weak place in the value chain.

    The money is going to Nvidia, who have the most powerful position.

    A bit like how a retailer can take all the margins of some innovative product, if they own the channel.

    AI is over-hyped, the Frontier Labs are over priced - but AI is here to stay, and will grow. Not like Skynet, but like a new form of compute. And it will take it's time, and the profits will be reaped by those with the power.

  5. randomImmigrant

    I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures.

    Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures.

    In narrow domains with more deterministic outputs though, this is less of an issue, and we see multiple agents succeed much better.

    The fusion of that capacity, with humans in the loop able to better direct such agents and act as their temporal tethers, is where I think the real action will be for a while at least.

  6. ausbah

    > the best alternative to rigorous specification is human review. human review doesn't scale well to the volumes of output produced by language models. to make matters worse

    when the business model is selling more tokens you get such per serve ice times that lead to “more” thinking, engagement baiting, fluffy narratives, and straight up dark patterns

  7. yunwal

    > those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc.

    I have no idea how people can so confidently say that call center work is a “controlled environment” or “repetitive”. It’s almost by definition not repetitive or controlled. Customer support is what I go to when the controlled environment has failed

  8. robinpie

    I really appreciate seeing a tempered take that's not literally denialist about current capabilities.

More from this day

2026-09-15