DeepSeek V4 Flash Analysis Fails Due to 404 Error

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

The anticipated analysis of DeepSeek V4 Flash regarding its intelligence, performance, and pricing is currently unavailable. Attempts to access the source URL from artificialanalysis.ai resulted in a 404 Not Found error. Consequently, no specific insights or data regarding the model's capabilities can be summarized at this time.

Target URL returned error 404: Not Found
  1. pmxi

    I have updated OpenAI's chart[1] from yesterday to include one more datapoint: DeepSeek V4 Flash 0731. It's on the frontier.

    https://files.parasmittal.com/openai_aa_luna_dsflash.svg

    1: https://openai.com/index/advancing-the-price-performance-fro...

  2. bwfan123

    > For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework

    So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.

  3. 0cf8612b2e1e

    Somewhat relatedly, how do the economics for Huggingface work? They must be hosting petabytes of models and datasets by now. I have downloaded quite a few “just in case”, only to replace them with the later iteration months later.

    Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?

  4. throwaw12

    If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?

  5. scosman

    So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon....

    Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.

  6. WithinReason

    Already beat Luna on price/task, by about 2x:

    https://artificialanalysis.ai/models/deepseek-v4-flash?intel...

  7. kamranjon

    The really interesting thing about this is how big of a jump was achieved with just extra fine-tuning here. No structural changes to the model, just more data, compute and time. It makes me pretty excited for the future of small models - DS v4 flash is a relatively small model when compared to the class it's competing with, so likely similar gains can be made applying quality data/training pipeline to other smaller models.

  8. ycui7

    For people with single RTX PRO 6000 96GB or DGX Spark 128GB, vllm-moet is a very good engine, although lesser known. It auto generate a symmetric 2-bit plane for inference and also generate a 4-bit delta cache to recover precision. Support ssd streaming oversized weight. You pick how much VRAM to allocate to each to balance out speed vs precision. 170 tps with ds-v4-flash demonstrated.

    It use the stock model, no new models requires.

    Worth spend a few hours to try.

    The DGX Spark requires a small hack to ignore the difference between sm120 vs sm121, but it does run on sm121.

  9. WhitneyLand

    It’s exciting that a model scoring this high is dirt cheap.

    It’s also so inefficient, when they release the full performance numbers it’s not going to be good.

    One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.

  10. coder543

    The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

More from this day

2026-07-31