DeepSeek V4 Flash official version launches, and the cheaper version actually beats the pricier one?
DeepSeek V4 Flash official version is live — and the cheap version actually beats the expensive one? At 7 AM this morning, DeepSeek posted on X, and the AI community went absolutely wild — the V4-Flash official API is now in public beta. Sounds like just a routine version update? But looking at the official benchmark chart, many industry insiders were stunned: 9 AI task capability metrics
💡 What You Will Learn
DeepSeek V4 Flash official version is live — and the cheap version actually beats the expensive one? At 7 AM this morning, DeepSeek posted on X, and the AI community went absolutely wild — the V4-Flas
📜 Table of Contents
DeepSeek V4 Flash Official Release — Did the Cheap Version Just Beat the Expensive One?
At 7 AM this morning, DeepSeek dropped a post on X, and the AI community went absolutely nuts—
The V4-Flash official API is now in public beta.
Sounds like just another routine version update, right? But when people saw the official benchmark chart, a lot of seasoned folks did a double take:
Across 9 AI capability tests, the cheap version outperformed the expensive preview version in every single one.
To put it simply: you buy an economy car, and it beats a sports model that costs three times more in the 0-100 km/h sprint.
This is worth digging into.
Bottom Line First: The Cheap Version Beat the Expensive One
DeepSeek's models come in two flavors:
- Flash: cheap, fast, positioned as "good enough"
- Pro: bigger, stronger, positioned as "built for heavy lifting"
By all logic, Pro should crush Flash across the board. But this time, the official release benchmarks tell a different story—Flash scored higher than the Pro preview in all 9 tests.
This isn't a "small improvement." This is "the economy model overtaking the luxury one."
What Actually Changed? The Answer: A "Renovation"
Don't jump to conclusions like "DeepSeek pulled off some new black magic."
Not a single parameter in the model's brain architecture was touched this update.
Same setup as before:
- 284 billion total parameters, but only 13 billion activated per query (technically called MoE — in plain terms, "a big team, but only the elite get deployed")
- Can read through a 1-million-character document in one go
- Three thinking modes: no thinking / deep thinking / max thinking
Identical to the previous preview version. So what changed?
What changed is "post-training" — think of it as "renovation."
Same building, but a bare concrete shell versus a fully renovated interior makes a world of difference in how it feels to live in. Same with models: the brain architecture (the shell) hasn't changed, but the training approach (the renovation) has been upgraded — and the results are completely different.
How Big Is the Score Gap? One Table Says It All
DeepSeek was generous this time and published all the test data. Let me translate it into plain English for you:
| Test Category | Cheap Flash | Expensive Pro Preview | Gap |
|---|---|---|---|
| Terminal Operations | 82.7 | 72.1 | +10.6 |
| Code Repository Reconstruction | 54.2 | 38.5 | +15.7 |
| Cybersecurity Hands-on | 76.7 | 52.7 | +24.0 |
| Software Development | 54.4 | 12.8 | +41.6 |
| Tool Use | 70.3 | 55.9 | +14.4 |
| Comprehensive Benchmark | 25.2 | 16.5 | +8.7 |
| Automated Testing | 25.1 | 12.8 | +12.3 |
| Full-Stack Development | 68.7 | 41.8 | +26.9 |
| Hardcore Development | 59.6 | 31.1 | +28.5 |
The most jaw-dropping row is "Software Development": the cheap version scored 54.4, while the expensive preview managed just 12.8 — a 41.6-point blowout.
What does that mean? It's like jumping from "barely passing" to "top of the class" on an exam — and the one who scored higher is the one who paid less.
Even more impressive: the cheap version's overall performance has already surpassed GLM 5.2, and in some categories it's approaching Claude Opus 4.8 — one of the strongest models out there.
Below is the benchmark chart DeepSeek officially posted this morning:
But Hold Your Applause — Three "Caveats" You Need to Know
Impressive scores, sure, but there are a few "buts" you should be aware of.
Caveat #1: These scores weren't achieved "naked."
The official team used their own in-house testing tool (called DeepSeek Harness), which hasn't been open-sourced yet. In other words, this is "taking the exam with the best equipment." If you run the model yourself, you may not reproduce these numbers.
Caveat #2: Two of the tests are "self-graded."
Full-Stack Development and Hardcore Development are from DeepSeek's internal question bank — outsiders can't verify them. Take the scores as a reference, but don't treat them as gospel.
Caveat #3: Real-world performance is what actually matters.
Official benchmarks are "lab conditions"; your projects are "survival in the wild." Between the two lies an entire real world.
What Does This Mean for You?
Two scenarios here.
If you're a regular user (chatting with ChatGPT or DeepSeek's web interface):
This update has nothing to do with you. It only changed the API. The mobile and web apps are untouched — the experience is identical to yesterday. Keep chatting, keep coding.
If you're a developer (using DeepSeek to run AI agents for projects):
Then there are three things worth serious thought—
First, it's time to re-evaluate "expensive vs. cheap."
Everyone defaulted to Pro before, because the cheap version couldn't handle complex tasks — it scored just 7.3 on "Software Development," which held a lot of people back. Now it's jumped to 54.4. Projects that previously chose Pro should be re-tested with the cheap version. You're not just saving money — you're saving multiple times the waiting time.
Second, official benchmarks are the "ceiling," not the "floor."
Those scores were achieved with the best equipment. If you get half of that in your own projects, that's already decent. But even half is better than the old Pro preview — worth the switch.
Third, the real big move might still be coming.
The official team said the testing tool (Harness) will be open-sourced soon, and the more expensive Pro official version is "coming soon" as well. If those two line up — that's the real game-changer.
As for the Pro official release, there was no update today, but the official "stay tuned" tease suggests it shouldn't be far off.
Uncle Peng's Take
One sentence to sum it up: Within this week, run one or two low-stakes tasks on the cheap version to test the waters.
If results are stable, then consider moving it to core projects. As for the Pro official version — its API updates in early August, and it'll likely ship with that open-source tool. For those waiting on Pro, this is what you've been waiting for.
One last honest thought: DeepSeek's move here shows me a trend — AI model competition has shifted from "who's bigger" to "who trains better." Same brain architecture, different training methods, wildly different results.
That's good news for everyday users: the more models compete, the lower the prices and the stronger the capabilities.
What score would you give DeepSeek for this move? Drop your thoughts in the comments👇
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
