Opus 5.5 runs coding tasks for hours with little oversight

Getting the most out of Opus 5.5 in Claude and Claude Code

Opus 5.5 runs coding tasks for hours with little oversight

Anthropic's guide to Opus 5.5 in Claude and Claude Code explains what changed: the model thinks before every reply, works longer on multi-part tasks, and reports what it did in plain language. The advice is to hand over the whole task, name the finish line, and delete "think carefully" prompts. Early testers had it run long coding jobs for hours, coordinate parallel subagents, and catch more bugs at its lowest effort than Opus 5 at high effort.

One early tester said Opus 5.5 at its lowest effort caught more bugs than Opus 5 at high effort, with fewer false alarms.
  1. rdli

    It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.

    9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.

  2. jjcm

    It’s extremely good at frontend, particularly if it has an image reference. I chucked in design reference images into this and told it to focus on the flowing svgs, and it crushed this Star Trek computer-inspired layout: https://html.non.io/lcars-opus-5.5

  3. pawelduda

    I pointed Opus 5.5 xhigh at a house construction blueprint (pdf with vector drawings) and asked it to create its 3D model in Blender. It one-shot the task in 45 min and outdid my (blender newbie) manual 50h+ work. It also flagged the same issues with the document I noticed earlier. I had some renders and plans from interior designer and then asked to blend them together. It nailed the job. $45 total API cost (I'm on a plan so it cost me way less, just posting what /usage shows). Crazy upgrade.

    I researched the feasibility of such task about a half year ago and concluded AI wouldn't be able to have good enough spatial and blueprint knowledge, unless you were willing to throw unreasonable amount of money at the task.

  4. adastra22

    Some of this advice is really missing the mark. I will speak to just one I know well. Many of my frequently used prompts have “think through this step by step” because if you don’t, it only considers the task holistically rather than step by step, and different issues emerge in that frame of thinking. I see this. Often when doing planning, for example, it will not notice interdependencies between tasks until you force it to think through doing the whole thing step by step (task by task) then it will notice that step 2 requires a feature introduced by step 14. It wouldn’t notice otherwise.

    Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.

    I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.

  5. hibikir

    It's much better than 5, but I've had a couple of situations this week where it was too interested in being independent, making calls that went directly against my recommendations. It can also do fun things like convince auto-mode to go way past what I have autorized. For instance, specific permission to run process X in region abz-1 suddenly became running X in 5 other regions, with no warning, and doing modifications that it never mentioned in the summaries. And a few of the times it got the calls very wrong, by assuming it understood systems it didn't. It'd even argue with me when corrected, as it assumed similar names were referring to the same thing, when they weren't.

    So asking it to do things on its own for a long time? Given last week, absolutely not.

  6. jampekka

    What is this spam of generic comments about how Opus 5.5 is so great, with perhaps some anecdote? How is that discussing the submission?

  7. magicalhippo

    Been very impressed with my most recent project. I wanted to simulate some older electronics circuits. I handed it a folder with scans of old service manuals which contained circuit diagrams. It managed to correctly interpret the circuits, including figuring out some were the same topology despite the diagrams being quite different, or some that had some subtle but very important differences despite looking almost identical at a glance.

    In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.

    It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.

    I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.

    It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.

    I also provided […]

  8. ToJans

    Superb model indeed.

    I've given it some big tasks and asked it to parallelize as much as possible etc.

    It did burn through my weekly tokens in about a day (20x max), but the output was completely on point.

    (I knew there was a "reset token usage - opus 5.5" button in my account.)

    I've now come to a point where I even delegate my discovery for new features to it.

    You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.

More from this day

2026-10-03