DeepSeek-V4-Flash Public Beta Launches with Enhanced Agent Capabilities
DeepSeek-V4-Flash Update

The DeepSeek-V4-Flash API is now in public beta, featuring significantly enhanced agent capabilities that outperform the V4-Pro-Preview across multiple benchmarks. This update natively supports the Responses API format for Codex integration while maintaining the same model architecture as the preview version. The release marks a major step forward in automated coding and tool usage, with the official V4-Pro version expected to follow soon.
The official release of the DeepSeek-V4-Flash API is now in public beta with significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview.
- NitpickLawyer
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.
DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.
Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.
- f311a
I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.
I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.
Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.
Also, it will never complain about security guards, I've been using it to reverse engineer binaries.
- lionkor
I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:
- Cost: $4.55USD
- API requests: 3,467
- Tokens: 323,183,886
And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks.
For everything else, use another model.
- kmarc
Essentially I'm running everything on flash now inside pi. With the correct set of MCP servers, context reducer tooling and skills it can implement any task I throw at it. Some sessions take 30+ turns, but it's fast and cheap; all this in an hour, with ~$0.5 cost.
(TBH though, in my multi-subagent workflow I do use other, more expensive models for planning, reviewing, oracle-ing)
I haven't used our slow opus subscription for weeks.
(Also set up an OpenWebUi self-hosted chat that works from my phone, has some mcp and skills. fully replaced perplexity. Monthly cost ~$18 for hosting and subscriptions)
- wkcheng
If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.
Crazy.
- dannyw
Deepseek and moonshot are the only two providers I consent to training for.
- ggcr
Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.
If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.
- heyalexej
Long story, I have humongous zai GLM 5.2 token budget that I'm using in a similar fashion as many comments explain here. GPT 5.6 or Fable 5 for planning, GLM for implementing, researching, extracting and many other tasks I consider grunt work. Very happy with the performance, speed isn't all that good though. I'd be curious to hear from someone who works with both, DeepSeek and GLM side by side.
- Goranek
Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?
Does this make sense?
- thirtygeo
For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself