Unreal Agent's async harness cuts coding agent costs by 40%

Unreal Agent's async harness cuts coding agent costs by 40%

Unreal Labs built Unreal Agent with a fully asynchronous tool-calling harness that removes waits, polls, and heartbeats from the model's workload. The result: up to 40% lower cost than Codex and 20% lower than Pi on real workloads and agentic benchmarks like Terminal-Bench 4.0, SWE-Atlas, and DeepSWE 1.1, with comparable pass rates. The SDK ships as a Go library, a runner executable, and a Harbor-compatible benchmark runner.

The Unreal Agent harness manages tool calls in a completely asynchronous way, relieving the underlying model of the need to manage waits, polls, and heartbeats for tools.
  1. dvt

    I think this space is very untapped. Models are interesting, but I am absolutely obsessed with some things I've been researching/working on for the past few years:

    Fractal tool discovery: tool taxonomy where an agent can "drill deeper" to find what specific tool it's looking for. Helps if/when polluting context with a zillion (mostly unnecessary) tools.

    Leveraging splay trees: this is my favorite data structure and I think relatively unused in the context of agents/harnesses. A lot of times, recently-used workflows/tool-chains will be used again, so having those at the top of the search hierarchy is an awesome optimization.

    Virtual containerized notebooks: models working in sandboxed (WASI) Python notebooks is incredible. Even local models (if given enough time) will usually converge on a good solution. Being able to mount tools/resources/fs is again, imo quite untapped. Some problems here are running native things (thing numpy/pandas) in containers is a nightmare (or impossible).

    Anyway, happy to see other folks seriously doing stuff in this space. If anyone wants to collaborate on anything don't hesitate to reach out :) I'm also actively looking for a job or some contract gigs.

    Fun times ahead.

  2. tekacs

    The headline graph is kind of bizarre.

    For some reason they're comparing their harness running on Astra xhigh to Codex with Astra max?

    ---

    Also worth noting that OpenAI just added support for async tool calling to their harness, which isn't 1:1 with this approach, but is slowly ramping up in being able to provide something similar.

    A big part of why Codex uses so many tokens is that it basically hot loops on polling tasks it starts for... absolutely no good reason: https://www.reddit.com/r/codex/comments/1wdlp7q/weve_discove...

    I fixed it on my fork of Codex too, also back in Jan/Feb – I keep this patch rebased, for anyone who wants it: https://github.com/tekacs/codex/commit/9ffcf8db9078eae43d411...

    It results in token savings similar in scale to those displayed here by Unreal.

    ---

    My harness has used a slightly fancier version of the approach that Unreal is using since ~Feb, and... it definitely works excellently, but it's also assuredly smoother with Astra and other recent models that are more aware of async tool calling.

  3. bryant

    So I commented on this in passing in a deeper thread (in re: the potential trademark issue - https://news.ycombinator.com/item?id=49807884), but I think this is serious enough for its own top-level thread.

    How likely is it that they get sued into the ground in a year? They might have a strong suite of offerings even as soon as six months from now, but if the essence of a company's brand seems at jeopardy from the start, can I take the risk as a potential customer that they'd survive that kind of action?

  4. tapoxi

    Sounds like a trademark issue when Epic ships a wildly popular Unreal Engine

  5. tontinton

    Oh very nice, can you also compare it to https://maki.sh?

    Would be interesting to compare to a harness optimizing for cost reduction too.

More from this day

2026-09-22