DeepSeek V4 Flash official version launches, and the cheaper version actually beats the pricier one?

📡 AI News 2026-08-27 5 min read

DeepSeek V4 Flash official version is live — and the cheap version actually beats the expensive one? At 7 AM this morning, DeepSeek posted on X, and the AI community went absolutely wild — the V4-Flash official API is now in public beta. Sounds like just a routine version update? But looking at the official benchmark chart, many industry insiders were stunned: 9 AI task capability metrics

💡 What You Will Learn

DeepSeek V4 Flash official version is live — and the cheap version actually beats the expensive one? At 7 AM this morning, DeepSeek posted on X, and the AI community went absolutely wild — the V4-Flas

📜 Table of Contents

DeepSeek V4 Flash Official Release — Did the Cheap Version Just Beat the Expensive One?

At 7 AM this morning, DeepSeek dropped a post on X, and the AI community went absolutely nuts—

The V4-Flash official API is now in public beta.

Sounds like just another routine version update, right? But when people saw the official benchmark chart, a lot of seasoned folks did a double take:

Across 9 AI capability tests, the cheap version outperformed the expensive preview version in every single one.

To put it simply: you buy an economy car, and it beats a sports model that costs three times more in the 0-100 km/h sprint.

This is worth digging into.

Bottom Line First: The Cheap Version Beat the Expensive One

DeepSeek's models come in two flavors:

By all logic, Pro should crush Flash across the board. But this time, the official release benchmarks tell a different story—Flash scored higher than the Pro preview in all 9 tests.

This isn't a "small improvement." This is "the economy model overtaking the luxury one."

What Actually Changed? The Answer: A "Renovation"

Don't jump to conclusions like "DeepSeek pulled off some new black magic."

Not a single parameter in the model's brain architecture was touched this update.

Same setup as before:

Identical to the previous preview version. So what changed?

What changed is "post-training" — think of it as "renovation."

Same building, but a bare concrete shell versus a fully renovated interior makes a world of difference in how it feels to live in. Same with models: the brain architecture (the shell) hasn't changed, but the training approach (the renovation) has been upgraded — and the results are completely different.

How Big Is the Score Gap? One Table Says It All

DeepSeek was generous this time and published all the test data. Let me translate it into plain English for you:

Test Category Cheap Flash Expensive Pro Preview Gap
Terminal Operations 82.7 72.1 +10.6
Code Repository Reconstruction 54.2 38.5 +15.7
Cybersecurity Hands-on 76.7 52.7 +24.0
Software Development 54.4 12.8 +41.6
Tool Use 70.3 55.9 +14.4
Comprehensive Benchmark 25.2 16.5 +8.7
Automated Testing 25.1 12.8 +12.3
Full-Stack Development 68.7 41.8 +26.9
Hardcore Development 59.6 31.1 +28.5

The most jaw-dropping row is "Software Development": the cheap version scored 54.4, while the expensive preview managed just 12.8 — a 41.6-point blowout.

What does that mean? It's like jumping from "barely passing" to "top of the class" on an exam — and the one who scored higher is the one who paid less.

Even more impressive: the cheap version's overall performance has already surpassed GLM 5.2, and in some categories it's approaching Claude Opus 4.8 — one of the strongest models out there.

Below is the benchmark chart DeepSeek officially posted this morning:

But Hold Your Applause — Three "Caveats" You Need to Know

Impressive scores, sure, but there are a few "buts" you should be aware of.

Caveat #1: These scores weren't achieved "naked."

The official team used their own in-house testing tool (called DeepSeek Harness), which hasn't been open-sourced yet. In other words, this is "taking the exam with the best equipment." If you run the model yourself, you may not reproduce these numbers.

Caveat #2: Two of the tests are "self-graded."

Full-Stack Development and Hardcore Development are from DeepSeek's internal question bank — outsiders can't verify them. Take the scores as a reference, but don't treat them as gospel.

Caveat #3: Real-world performance is what actually matters.

Official benchmarks are "lab conditions"; your projects are "survival in the wild." Between the two lies an entire real world.

What Does This Mean for You?

Two scenarios here.

If you're a regular user (chatting with ChatGPT or DeepSeek's web interface):

This update has nothing to do with you. It only changed the API. The mobile and web apps are untouched — the experience is identical to yesterday. Keep chatting, keep coding.

If you're a developer (using DeepSeek to run AI agents for projects):

Then there are three things worth serious thought—

First, it's time to re-evaluate "expensive vs. cheap."

Everyone defaulted to Pro before, because the cheap version couldn't handle complex tasks — it scored just 7.3 on "Software Development," which held a lot of people back. Now it's jumped to 54.4. Projects that previously chose Pro should be re-tested with the cheap version. You're not just saving money — you're saving multiple times the waiting time.

Second, official benchmarks are the "ceiling," not the "floor."

Those scores were achieved with the best equipment. If you get half of that in your own projects, that's already decent. But even half is better than the old Pro preview — worth the switch.

Third, the real big move might still be coming.

The official team said the testing tool (Harness) will be open-sourced soon, and the more expensive Pro official version is "coming soon" as well. If those two line up — that's the real game-changer.

As for the Pro official release, there was no update today, but the official "stay tuned" tease suggests it shouldn't be far off.

Uncle Peng's Take

One sentence to sum it up: Within this week, run one or two low-stakes tasks on the cheap version to test the waters.

If results are stable, then consider moving it to core projects. As for the Pro official version — its API updates in early August, and it'll likely ship with that open-source tool. For those waiting on Pro, this is what you've been waiting for.


One last honest thought: DeepSeek's move here shows me a trend — AI model competition has shifted from "who's bigger" to "who trains better." Same brain architecture, different training methods, wildly different results.

That's good news for everyday users: the more models compete, the lower the prices and the stronger the capabilities.

What score would you give DeepSeek for this move? Drop your thoughts in the comments👇

Related Articles
2026-08-23
Hermes Agent Gets a "Slimming" Update: Database Compression Up to 78%, I Tested 600MB Cut to 270MB
2026-09-14
Memory prices have surged this much, and AMD has dug out a five-year-old socket to sell: the $99 Ryzen 5500F is here
2026-10-10
OpenAI's math proof storm challenges academia, but mathematicians aren't buying it
2026-07-09
OpenAI's free tier is now a completely different product
2026-08-01
Grok 4.5 Has Been Running at SpaceX for a Week—You Need to Read Musk's Tweet Carefully
2026-07-13
AMD RDNA 5 GPUs Not Coming Until Late 2027 at the Earliest — Computex Leaks Reveal

Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.

💬 Comments (0)

No comments yet. Be the first!

Login to comment